SurfMind Logosurfmind
  • Chat
  • Mô hình
  • Kỹ năng
  • Bảng giá
  • Tải xuống
chromeThêm vào Chrome
SurfMind Logo
surfmind

Trợ lý AI cho mọi trang web bạn truy cập

wave@surfmind.ai

Ứng dụng

  • Tiện ích Chrome
  • Tiện ích Firefox
  • Tiện ích Safari
  • Tiện ích Safari iOS

Tài nguyên

  • Tài khoản
  • Bảng giá
  • Có gì mới
  • Cách thiết lập

Khám phá

  • Bài viết
  • Mô hình
  • Bảng xếp hạng
  • So sánh mô hình

Công ty

  • Chương trình tiếp thị liên kết
  • Chính sách quyền riêng tư
  • Điều khoản sử dụng
© 2026 SurfMind. Bảo lưu mọi quyền.

Mô hình

Khám phá mọi mô hình AI có trên SurfMind và so sánh giá, cửa sổ ngữ cảnh cùng tốc độ.

So sánhBảng xếp hạng
  • GPT-6 Sol vs Claude Opus 5.5
  • Claude Opus 5.5 vs Claude Sonnet 5.5
  • Claude Sonnet 5.5 vs GPT-5.6 Terra
  • Grok 4.6 vs GPT-6 Sol
  • Claude Opus 5.5 vs Gemini 3.1 Pro Preview
  • DeepSeek V3.2 vs GPT-6 Sol
  • Gemini 3.1 Pro Preview vs GPT-6 Sol
  • DeepSeek V3.2 vs Claude Opus 5.5
  • Claude Haiku 4.5 vs GPT-6 Luna
  • Llama 4 Maverick vs DeepSeek V3.2
  • Mistral Large vs Llama 4 Maverick
  • Kimi K3 vs Claude Fable 5
  • GLM 5.2 vs DeepSeek V4 Pro 0813
  • Ling 3.0 Flash vs DeepSeek V4 Flash 0731
⌘K

450 mô hình

  • Mô hình

    Ling 3.1 Flash

    InclusionAI🇨🇳

    Ling 3.1 Flash is InclusionAI's (Ant Group) hybrid-reasoning Mixture-of-Experts model (560B total, 25B active parameters, 262K-token context), a much larger successor to the 124B Ling 3.0 Flash. This OpenRouter listing is free of token cost, making it a good fit for trying out strong reasoning, coding, and agentic workloads without spend.

    Oct 2, 2026·ngữ cảnh 262k·Miễn phí
  • Apodex 1.1 Mini (free)

    apodex

    This is the free OpenRouter endpoint for Apodex 1.1 Mini, a reasoning-first, Apache-2.0 mixture-of-experts model (about 35B total, 3B active parameters, 262K-token context) with native function calling. It is built for long-horizon research and forecasting that works directly with files, data, code, and tools to produce verifiable results; free access is rate-limited, so it suits experimentation over production use.

    Oct 1, 2026·ngữ cảnh 262k·Miễn phí
  • Pareto 26.10 Preview

    unbiased

    Pareto 26.10 Preview is a preview build of Unbiased's proprietary multimodal composite model (text and image input, 1M-token context, up to 128K output tokens), built for research, coding, and agentic workflows. It targets frontier-level performance on general-purpose tasks at mid-tier pricing ($0.80/M input, $3.20/M output), so expect behavior to change before the next stable Pareto release.

    Oct 1, 2026·ngữ cảnh 1M·mức chi phí 0.7x
  • Mô hình

    GPT-6.1 Sol Pro

    OpenAI🇺🇸

    GPT-6.1 Sol Pro is OpenAI's GPT-6.1 Sol served with reasoning mode set to "pro" for higher-quality answers on complex tasks, keeping the same 1.05M-token context and $2/$10 per-million input/output pricing. It suits difficult coding, analysis, and research problems where extra reasoning is worth more tokens and latency.

    Sep 29, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT-6.1 Sol

    OpenAI🇺🇸
    Pháp lý & Chính phủ#4Phần mềm & Dịch vụ CNTT#5Hiểu hình ảnh#10Giải trí, Thể thao & Truyền thông#16

    GPT-6.1 Sol is OpenAI's upgrade to GPT-6 Sol (1.05M-token context, $2/$10 per million input/output tokens), released September 29, 2026 at DevDay below the flagship GPT-6 Astra. OpenAI says it nearly matches Astra on agentic coding, computer use, and professional work at a fifth of the price, making it a strong default for coding agents and document-heavy workflows.

    Sep 29, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    Claude Sonnet 5.5

    Anthropic🇺🇸
    Phần mềm & Dịch vụ CNTT#20Viết, Văn học & Ngôn ngữ#34Pháp lý & Chính phủ#35Hiểu hình ảnh#35

    Claude Sonnet 5.5 is Anthropic's Sonnet-class model (1M-token context, 128K max output, $2/$10 per million input/output tokens), released September 28, 2026 as a direct upgrade to Sonnet 5 and a faster, lower-cost complement to Opus 5.5. It generates output over 30% faster with fewer tokens per task, and is strongest at well-scoped everyday work such as building features, fixing bugs, and producing polished documents, slides, and spreadsheets.

    Sep 28, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Claude Sonnet 5.5 (batch)

    Anthropic🇺🇸
    Phần mềm & Dịch vụ CNTT#20Viết, Văn học & Ngôn ngữ#34Pháp lý & Chính phủ#35Hiểu hình ảnh#35

    Batch-processing variant of Claude Sonnet 5.5, Anthropic's fast Sonnet-class model for everyday coding and document work. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Sep 28, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    Perceptron Mk1.5

    Perceptron🇺🇸

    Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents (36K-token context, $0.15/$1.50 per million input/output tokens), released September 25, 2026. It takes text, image, video, and audio input and answers with text plus optional structured annotations (points, boxes, polygons, tracks, and clips), adding audio, video tracking, tool calling, and graded reasoning over Mk1, which makes it a fit for robots, drones, smart glasses, and other perception-driven agents.

    Sep 25, 2026·ngữ cảnh 37k·mức chi phí 0.3x
  • Mô hình

    Ember-1

    Fireworks🇺🇸

    Ember-1 is Fireworks Research's reasoning model built on Moonshot AI's Kimi K3, retrained (not a new base model) to produce reasoning traces roughly 40% shorter while keeping K3-level quality, with text and image input and a 1M-token context window. Released September 23, 2026 as a research preview, it targets coding, knowledge work, and multi-turn agent workflows where reasoning tokens drive most of the cost and latency.

    Sep 24, 2026·ngữ cảnh 1M·mức chi phí 3x
  • Mô hình

    GLM 5.3 Prime

    Z.AI🇨🇳

    GLM-5.3-Prime is the high-speed serving tier of Z.ai's open-weight GLM-5.3 flagship, running the same weights (1M-token context, up to 128K output) behind an accelerated stack for 1.5-2x the output throughput at $2.80/$8.80 per million input/output tokens. Reasoning is always on (low, high, or max effort), and it targets latency-sensitive coding, streaming code generation, and long-horizon multi-turn agent orchestration.

    Sep 23, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Qwen3.8 Max Prime

    Qwen🇨🇳

    Qwen3.8 Max Prime is the high-speed serving tier of Alibaba's proprietary Qwen3.8 Max (same 2.4-trillion-parameter MoE weights, 1M-token context, text, image, and video input, reasoning on by default), delivering 1.5-2x the output throughput at a higher price of $4/$12 per million input/output tokens. Announced September 22, 2026 as Alibaba's "Prime mode", it is for coding, office automation, and long-running agent workflows where latency matters more than cost.

    Sep 23, 2026·ngữ cảnh 1M·mức chi phí 2.5x
  • Mô hình

    Aion 3.5 Mini

    Aion Labs🇵🇱

    Aion 3.5 Mini is the smaller, lower-cost sibling of AionLabs' Aion 3.5, a GLM-based multi-model roleplaying and storytelling system with a 262K-token context window. It uses the same collaborative generation process as Aion 3.5 at roughly a quarter of the price, suited to high-volume or latency-sensitive roleplay and creative writing.

    Sep 23, 2026·ngữ cảnh 262k·mức chi phí 0.3x
  • Mô hình

    Aion 3.5

    Aion Labs🇵🇱

    Aion 3.5 is AionLabs' multi-model roleplaying and storytelling system built on the GLM family (262K-token context, released September 23, 2026), in which several specialized models each contribute to every response. It succeeds Aion-3.0 and is tuned for character-driven fiction with stronger narrative structure, tension, and conflict, rather than coding or general assistant work.

    Sep 23, 2026·ngữ cảnh 262k·mức chi phí 1.5x
  • Mô hình

    Solar Mini 4

    Upstage🇰🇷

    Solar Mini 4 is Upstage's compact mixture-of-experts model (35B total, 3B active parameters, 524K-token context), released September 23, 2026 at $0.05/$0.20 per million input/output tokens. It is built for fast, low-cost agentic use with tool calling and structured outputs, with fluent Korean alongside strong English and Japanese, as a cheaper sibling to Solar Pro 4.

    Sep 23, 2026·ngữ cảnh 524k·mức chi phí 0.1x
  • Mô hình

    Command A+

    Cohere🇨🇦

    Command A+ is Cohere's flagship enterprise model, an open-weight sparse mixture-of-experts (218B total, 25B active parameters, Apache 2.0) with text and image input, 128K-token input context, and support for 48 languages, released May 20, 2026. It combines reasoning, RAG with native citations, coding, and strict tool calling in one model that runs on as few as two H100 GPUs, aimed at privately deployable agentic workflows.

    Sep 22, 2026·ngữ cảnh 192k·mức chi phí 0.3x
  • Mô hình

    GPT-6 Luna Pro

    OpenAI🇺🇸

    GPT-6 Luna Pro is OpenAI's GPT-6 Luna served with reasoning mode set to "pro" for higher-quality answers on harder problems, keeping the same 1.05M-token context and $0.10/$0.50 per-million input/output pricing as the base model. It suits workloads that want Luna's low cost but need more careful reasoning than the default mode, at the expense of more tokens and latency per request.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    GPT-6 Luna Pro (batch)

    OpenAI🇺🇸

    This is the batch-processing variant of GPT-6 Luna Pro, OpenAI's low-cost GPT-6 Luna served with reasoning mode "pro", offered at half the standard token price. Requests run asynchronously, best suited to large-scale jobs that benefit from deeper reasoning but do not need real-time responses.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    GPT-6 Luna

    OpenAI🇺🇸
    Toán học#60Phần mềm & Dịch vụ CNTT#64Giải trí, Thể thao & Truyền thông#76Kinh doanh, Quản lý & Tài chính#82

    GPT-6 Luna is the fast, low-cost tier of OpenAI's GPT-6 series (1.05M-token context, text, image, and file input, $0.10/$0.50 per million input/output tokens), released September 22, 2026 below GPT-6 Sol and the flagship GPT-6 Astra. It is built for high-volume tasks with a clear goal, such as summarization, extraction, classification, quick questions, and lightweight agentic steps.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    GPT-6 Luna (batch)

    OpenAI🇺🇸
    Toán học#60Phần mềm & Dịch vụ CNTT#64Giải trí, Thể thao & Truyền thông#76Kinh doanh, Quản lý & Tài chính#82

    This is the batch-processing variant of GPT-6 Luna, the fast, low-cost tier of OpenAI's GPT-6 series with a 1.05M-token context window, offered at half the standard token price. Requests run asynchronously, making it a good fit for very large summarization, extraction, and classification jobs that do not need real-time responses.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    GPT-6 Sol Pro

    OpenAI🇺🇸

    GPT-6 Sol Pro is OpenAI's GPT-6 Sol served with reasoning mode set to "pro" for higher-quality answers on complex tasks, keeping the same 1.05M-token context and $2/$10 per-million input/output pricing as the base model. It targets difficult coding, analysis, and research problems where extra reasoning is worth more tokens and latency but Astra's price is not justified.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT-6 Sol Pro (batch)

    OpenAI🇺🇸

    This is the batch-processing variant of GPT-6 Sol Pro, OpenAI's GPT-6 Sol served with reasoning mode "pro" for maximum quality, offered at half the standard token price. Requests run asynchronously, suited to large-scale jobs that need deeper reasoning but not real-time responses.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 1x
  • Mô hình

    GPT-6 Sol

    OpenAI🇺🇸
    Phần mềm & Dịch vụ CNTT#44Giải trí, Thể thao & Truyền thông#49Viết, Văn học & Ngôn ngữ#71Khoa học Sự sống, Vật lý & Xã hội#71

    GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series (1.05M-token context, text, image, and file input, $2/$10 per million input/output tokens), released September 22, 2026 between the flagship GPT-6 Astra and the fast GPT-6 Luna. OpenAI reports it makes about half as many factual mistakes as GPT-5.6 Sol, reaching Astra-level reliability for everyday agentic coding and professional work at much lower cost.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT-6 Sol (batch)

    OpenAI🇺🇸
    Phần mềm & Dịch vụ CNTT#44Giải trí, Thể thao & Truyền thông#49Viết, Văn học & Ngôn ngữ#71Khoa học Sự sống, Vật lý & Xã hội#71

    This is the batch-processing variant of GPT-6 Sol, OpenAI's cost-efficient high-end GPT-6 model with a 1.05M-token context window, offered at half the standard token price. Requests run asynchronously, making it best for large coding, analysis, and document jobs that do not require real-time responses.

    Sep 22, 2026·ngữ cảnh 1.1M·mức chi phí 1x
  • Mô hình

    Claude Opus 5.5

    Anthropic🇺🇸
    Viết, Văn học & Ngôn ngữ#2Giải trí, Thể thao & Truyền thông#2Y học & Chăm sóc sức khỏe#2Toán học#6

    Claude Opus 5.5 is Anthropic's flagship model released September 22, 2026 as the first of the Claude 5.5 family, with a 1M-token context window, text, image, and file input, and pricing of $4/$20 per million input/output tokens. Anthropic says it performs at the level of Claude Fable 5.1 on most work while costing about 40% less to run than Opus 5, and it leads on agentic coding, multi-step changes in large codebases, and long-horizon knowledge work.

    Sep 22, 2026·ngữ cảnh 1M·mức chi phí 4x
  • Mô hình

    Claude Opus 5.5 (batch)

    Anthropic🇺🇸
    Viết, Văn học & Ngôn ngữ#2Giải trí, Thể thao & Truyền thông#2Y học & Chăm sóc sức khỏe#2Toán học#6

    Batch-processing variant of Claude Opus 5.5, Anthropic's Claude 5.5 flagship for agentic coding and long-horizon knowledge work with a 1M-token context window. Same capabilities as the standard model, offered at half the token price via asynchronous batch processing for large jobs that do not need real-time responses.

    Sep 22, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    MiMo-V2.6-Pro-UltraSpeed

    Xiaomi🇨🇳

    MiMo-V2.6-Pro-UltraSpeed is a fast-serving edition of Xiaomi's flagship MiMo-V2.6-Pro (MoE, over 1 trillion parameters, 1M-token context, omnimodal text/image/audio/video input), running the same checkpoint at roughly 10x the output speed of the standard endpoint with matching quality. It is built for latency-sensitive interactive agents that still need Pro-level coding, reasoning, and multimodal capability.

    Sep 21, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    MiMo-V2.6-Flash

    Xiaomi🇨🇳
    Phần mềm & Dịch vụ CNTT#38Hiểu hình ảnh#53Kinh doanh, Quản lý & Tài chính#64Toán học#70

    MiMo-V2.6-Flash is Xiaomi's open-source foundation model (MoE, 309B total parameters, 15B active per token, 1M-token context, omnimodal text/image/audio/video input), a lower-cost, hybrid-attention sibling to the closed flagship MiMo-V2.6-Pro. It targets high-volume workloads that need most of Pro's agentic coding and reasoning capability at a fraction of the inference cost.

    Sep 21, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    MiMo-V2.6-Pro

    Xiaomi🇨🇳
    Phần mềm & Dịch vụ CNTT#10Toán học#16Khoa học Sự sống, Vật lý & Xã hội#30Giải trí, Thể thao & Truyền thông#37

    MiMo-V2.6-Pro is Xiaomi's flagship foundation model (MoE, over 1 trillion parameters, 1M-token context, omnimodal text/image/audio/video input), built to push the ceiling of capability in Xiaomi's model lineup. It targets the most demanding coding, reasoning, and agentic workloads, with a faster UltraSpeed edition and a cheaper, open-weight Flash sibling available for less latency- or cost-sensitive use.

    Sep 21, 2026·ngữ cảnh 1.1M·mức chi phí 0.2x
  • Mô hình

    Grok 4.7

    xAI🇺🇸
    Pháp lý & Chính phủ#72Y học & Chăm sóc sức khỏe#72Viết, Văn học & Ngôn ngữ#80Giải trí, Thể thao & Truyền thông#85

    Grok 4.7 is xAI's (listed on OpenRouter under the SpaceXAI brand) proprietary flagship model, succeeding Grok 4.6, with a 500K-token context window and up to 450K output tokens. It is positioned for long-running software-engineering agents and knowledge work, with particular emphasis on verifying its own work over extended, multi-step tasks.

    Sep 21, 2026·ngữ cảnh 500k·mức chi phí 1x
  • Mô hình

    Qwen3.8 Omni Flash

    Qwen🇨🇳

    Qwen3.8 Omni Flash is Alibaba's omni-modal reasoning model (text, image, audio, and video input, 1M-token context, thinking on by default), released September 18, 2026 as a hosted API with no open weights. It is the first Qwen model built around agentic audio-video work such as video editing, audio-visual summarization, and real-time conversation, with speech recognition in 74 languages and very low audio input pricing.

    Sep 21, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Ternary Bonsai 2 27B

    prism-ml

    Ternary Bonsai 2 27B is PrismML's ternary-quantized derivative of Qwen3.8-27B (27B dense parameters, Apache 2.0, 256K-token context), compressing weights to roughly 1.7 bits per parameter so the full model fits in under 9 GB while retaining most of the original's benchmark score. It targets coding, math, tool-calling, and image-understanding workloads that need to run locally on a single consumer GPU or high-end laptop rather than in the cloud.

    Sep 18, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GLM 5.3 FlashX

    Z.AI🇨🇳

    GLM-5.3-FlashX is Z.ai's speed-optimized variant of GLM-5.3-Flash, a natively multimodal MoE model (around 320B total / 18B active parameters, MIT-licensed, 1M-token context) built on the same hybrid sparse-and-linear attention architecture. It generates at up to 200 tokens/sec for latency-sensitive coding and long-horizon agent tasks that still need the full 1M-token context, at a premium over the standard GLM-5.3-Flash endpoint.

    Sep 18, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Pareto

    unbiased

    Pareto is Unbiased's proprietary multimodal model (256K-token context, up to 128K output tokens), described by its maker as a composite system built for research, coding, and agentic workflows. It aims for frontier-level performance across general-purpose tasks, priced in the mid tier between fast, cheap models and top frontier flagships.

    Sep 17, 2026·ngữ cảnh 262k·mức chi phí 1.5x
  • Mô hình

    DeepSeek Pro Latest

    DeepSeek🇨🇳
    Pháp lý & Chính phủ#37Kinh doanh, Quản lý & Tài chính#44Viết, Văn học & Ngôn ngữ#47Khoa học Sự sống, Vật lý & Xã hội#48

    This is OpenRouter's auto-updating alias to DeepSeek's current top-tier 'Pro' model, always resolving to the newest release in that family (1M-token context) rather than a pinned snapshot. It lets integrations target DeepSeek's flagship reasoning line by a stable identifier without updating the model string on every new checkpoint.

    Sep 14, 2026·ngữ cảnh 1M·mức chi phí 0.7x
  • Mô hình

    DeepSeek Flash Latest

    DeepSeek🇨🇳
    Toán học#15Y học & Chăm sóc sức khỏe#16Phần mềm & Dịch vụ CNTT#22Kinh doanh, Quản lý & Tài chính#27

    This is OpenRouter's auto-updating alias to DeepSeek's current 'Flash' model, always resolving to the newest, lower-cost release in that family (1M-token context) rather than a pinned snapshot. It lets integrations track DeepSeek's newest fast, cheap model line, currently the DeepSeek V4.1 Flash generation, without updating the model string on every release.

    Sep 14, 2026·ngữ cảnh 1M·mức chi phí 0.4x
  • Schematron V2 Turbo

    inference-net

    Schematron V2 Turbo is Inference.net's 3B-parameter HTML-to-JSON extraction model (128K-token context), built on an IBM Granite 4.0 H Micro base and tuned for throughput over quality in high-volume scraping workloads. Extraction schemas are supplied via response_format rather than prompts, and it trades a little accuracy for speed against its sibling, Schematron V2 Small.

    Sep 12, 2026·ngữ cảnh 128k·mức chi phí 0.1x
  • Schematron V2 Small

    inference-net

    Schematron V2 Small is Inference.net's 3B-parameter HTML-to-JSON extraction model (128K-token context), built on a Llama 3.2 3B base and tuned for extraction quality on complex schemas and long pages rather than raw throughput. Like its Turbo sibling, it takes the target schema through response_format, making it a fit for accuracy-sensitive scraping and data-pipeline work.

    Sep 12, 2026·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    GPT Astra Latest

    OpenAI🇺🇸
    Phân tích tài liệu#16Phần mềm & Dịch vụ CNTT#17Hiểu hình ảnh#20Giải trí, Thể thao & Truyền thông#22

    This is OpenRouter's auto-updating alias to OpenAI's current 'Astra' model, always resolving to the newest release in that family (1.05M-token context) rather than a pinned snapshot. Astra is priced well above Sol, Terra, and Luna in the same alias family, positioning it as OpenAI's most expensive and most capable current tier on OpenRouter.

    Sep 11, 2026·ngữ cảnh 1.1M·mức chi phí 10x
  • Mô hình

    GPT Sol Latest

    OpenAI🇺🇸
    Pháp lý & Chính phủ#4Phần mềm & Dịch vụ CNTT#5Hiểu hình ảnh#10Giải trí, Thể thao & Truyền thông#16

    This is OpenRouter's auto-updating alias to OpenAI's current 'Sol' model, always resolving to the newest release in that family (1.05M-token context) rather than a pinned snapshot. Sol has been positioned as OpenAI's flagship-tier alias, priced below the newer, more expensive Astra alias in the same family.

    Sep 11, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT Terra Latest

    OpenAI🇺🇸
    Phân tích tài liệu#13Toán học#33Hiểu hình ảnh#36Pháp lý & Chính phủ#43

    This is OpenRouter's auto-updating alias to OpenAI's current 'Terra' model, always resolving to the newest release in that family (1.05M-token context) rather than a pinned snapshot. Terra sits in the mid tier by price among OpenAI's Sol/Terra/Luna/Astra alias family, between the cheaper Luna and the pricier Sol and Astra.

    Sep 11, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT Luna Latest

    OpenAI🇺🇸
    Toán học#60Phần mềm & Dịch vụ CNTT#64Giải trí, Thể thao & Truyền thông#76Kinh doanh, Quản lý & Tài chính#82

    This is OpenRouter's auto-updating alias to OpenAI's current 'Luna' model, always resolving to the newest release in that family (1.05M-token context) rather than a pinned snapshot. Luna is the cheapest and fastest tier in OpenAI's Sol/Terra/Luna/Astra alias family, priced roughly ten times below Sol and suited to high-volume, latency-sensitive chat and lightweight agent work.

    Sep 11, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    Fugu Ultra v2

    Sakana AI🇯🇵

    Fugu Ultra v2 is the higher-performance tier of Sakana AI's Fugu family: rather than a single trained checkpoint, OpenRouter describes it as a learned multi-agent orchestration system that routes each request across a pool of models, exposed here with a 1M-token context. It is proprietary and priced at the top of the Fugu line, aimed at the hardest coding, reasoning, and tool-use tasks where Fugu Max's lower cost is not enough.

    Sep 11, 2026·ngữ cảnh 1M·mức chi phí 6x
  • Mô hình

    Fugu Max

    Sakana AI🇯🇵

    Fugu Max is the cost-performance tier of Sakana AI's Fugu family: like Fugu Ultra v2, OpenRouter describes it as a learned multi-agent orchestration system rather than a single checkpoint, routing requests across models behind a 1M-token-context endpoint. It is proprietary and priced well below Fugu Ultra v2, fitting routine and moderately difficult coding and interactive tasks.

    Sep 11, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    Ling 3.0 Flash VL

    InclusionAI🇨🇳

    Ling 3.0 Flash VL is InclusionAI's open-weight multimodal MoE model (124B total / 5.5B active parameters, 128K-token context on this metered endpoint), extending the text-only Ling 3.0 Flash with native visual perception for image and video input. It is aimed at vision-grounded agentic tasks that route image and video understanding through the same reasoning pipeline used for text.

    Sep 10, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    DeepSeek V4.1 Flash

    DeepSeek🇨🇳
    Toán học#15Y học & Chăm sóc sức khỏe#16Phần mềm & Dịch vụ CNTT#22Kinh doanh, Quản lý & Tài chính#27

    DeepSeek V4.1 Flash is DeepSeek's open-weight sparse MoE model and the first built on the company's Causal Encoder-Decoder architecture, activating 8B parameters on input and 16B on output with a 1M-token context window. It targets coding, terminal and computer-use agents, and long-horizon multi-step tasks as a low-cost, high-throughput sibling to DeepSeek's Pro line.

    Sep 10, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    DeepSeek V4.1 Flash (batch)

    DeepSeek🇨🇳
    Toán học#15Y học & Chăm sóc sức khỏe#16Phần mềm & Dịch vụ CNTT#22Kinh doanh, Quản lý & Tài chính#27

    Batch-processing variant of DeepSeek V4.1 Flash, DeepSeek's open-weight sparse MoE model built on its Causal Encoder-Decoder architecture (8B active parameters on input, 16B on output, 1M-token context). Same coding and agentic capabilities as the standard model, offered at lower cost via asynchronous batch processing for high-volume jobs that do not need real-time responses.

    Sep 10, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Mercury 2.5

    Inception🇺🇸

    Mercury 2.5 is Inception Labs' diffusion-based LLM (dLLM), generating tokens in parallel rather than sequentially to reach about 1,107 tokens/second with a 260K-token context window, and is billed as the largest diffusion language model trained to date. It delivers a roughly 40% intelligence gain over Mercury 2 (comparable to cost-optimized frontier models like Claude Haiku 4.5, Gemini 3.5 Flash-Lite, and GPT-5.6 Luna Low) with tunable reasoning, parallel tool calls, and schema-aligned JSON, suiting latency-sensitive search agents, voice pipelines, and coding subagents.

    Sep 8, 2026·ngữ cảnh 260k·mức chi phí 0.1x
  • Mô hình

    Nex-N2.5-Mini

    Nex AGI🇨🇳

    Nex-N2.5-Mini is the smallest tier of Nex AGI's September 2026 Nex-N2.5 family, a multimodal agentic mixture-of-experts model (about 35B total parameters, 262K-token context) released with Apache-2.0 open weights. It targets computer use, browser automation, and agentic coding at very low per-token cost, making it a cheap default for high-volume agent loops.

    Sep 8, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Nex-N2.5-Pro

    Nex AGI🇨🇳

    Nex-N2.5-Pro is Nex AGI's multimodal agentic mixture-of-experts model (about 397B total parameters, 262K-token context), the middle tier of the September 2026 Nex-N2.5 family above Mini. It builds on Nex-N2 with focused gains in computer use, web browsing, and visually grounded agentic coding, suiting long-horizon agent workloads that need a visual feedback loop.

    Sep 8, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GPT-6 Astra

    OpenAI🇺🇸
    Phân tích tài liệu#16Phần mềm & Dịch vụ CNTT#17Hiểu hình ảnh#20Giải trí, Thể thao & Truyền thông#22

    GPT-6 Astra is OpenAI's flagship GPT-6-generation model (1.05M-token context, up to 128K output tokens, priced at $10/$50 per million input/output tokens), built for demanding end-to-end work. It excels at advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strength in long-horizon agentic tasks involving computer and browser use.

    Sep 4, 2026·ngữ cảnh 1.1M·mức chi phí 10x
  • Mô hình

    GPT-6 Astra (batch)

    OpenAI🇺🇸
    Phân tích tài liệu#16Phần mềm & Dịch vụ CNTT#17Hiểu hình ảnh#20Giải trí, Thể thao & Truyền thông#22

    This is the batch-processing variant of OpenAI's flagship GPT-6 Astra, offering the same 1.05M-token context and long-horizon agentic, coding, and research capabilities at lower cost for asynchronous workloads. Requests run on a delayed, non-interactive basis, making it best for large-scale jobs that do not require real-time responses.

    Sep 4, 2026·ngữ cảnh 1.1M·mức chi phí 5x
  • Mô hình

    GPT-6 Astra Pro

    OpenAI🇺🇸

    GPT-6 Astra Pro is OpenAI's flagship GPT-6 Astra served with reasoning mode set to "pro" for higher-quality answers on the hardest tasks, keeping the same 1.05M-token context window and $10/$50 per-million input/output pricing as the base model. It targets demanding end-to-end work such as advanced analysis, software engineering, deep research, scientific work, and long-horizon agentic tasks involving computer and browser use, trading extra latency and reasoning cost for maximum quality.

    Sep 4, 2026·ngữ cảnh 1.1M·mức chi phí 10x
  • Mô hình

    GPT-6 Astra Pro (batch)

    OpenAI🇺🇸

    This is the batch-processing variant of GPT-6 Astra Pro, OpenAI's flagship GPT-6 Astra served with reasoning mode "pro" for maximum quality, offered at lower cost for asynchronous workloads. Requests run on a delayed, non-interactive basis, best suited to large-scale jobs that do not need real-time responses.

    Sep 4, 2026·ngữ cảnh 1.1M·mức chi phí 5x
  • Mô hình

    Ling 3.0 Flash Sante (free)

    InclusionAI🇨🇳

    This is the free OpenRouter endpoint for InclusionAI's (Ant Group) Ling-3.0-Flash-Sante, a health and medicine fine-tune of the Ling 3.0 Flash base (124B total, about 5.1B active parameters, 262K-token context), released September 4, 2026. It targets medical knowledge reasoning, clinical safety, and evidence-based retrieval (leading its size class on benchmarks like MedXpertQA-Text and DiagnosisArena-MCQ) while retaining general reasoning and coding ability; free access is rate-limited and best suited for evaluation rather than production.

    Sep 4, 2026·ngữ cảnh 262k·Miễn phí
  • Mô hình

    Qwen3.8 Max (0902)

    Qwen🇨🇳
    Hiểu hình ảnh#2Khoa học Sự sống, Vật lý & Xã hội#11Y học & Chăm sóc sức khỏe#12Viết, Văn học & Ngôn ngữ#21

    Qwen3.8-Max-0902 is Alibaba's September 2, 2026 snapshot of its proprietary flagship Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input and has reasoning on by default. The architecture and pricing are unchanged from the base model; this post-training update sharpens coding and agentic performance across multi-step software projects, multi-tool orchestration, and long-horizon task execution.

    Sep 3, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    Muse Spark 1.3 Contributor

    Meta🇺🇸

    This is the discounted 'contributor' tier of Meta's Muse Spark 1.3 coding and agentic model, priced far below the standard endpoint (about $0.10/$0.20 per million input/output tokens versus $1.25/$4.25) in exchange for letting Meta use your prompts and completions for training. Capabilities match the standard model, so it is best reserved for non-sensitive workloads where that data-sharing tradeoff is acceptable.

    Sep 2, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Muse Spark 1.3

    Meta🇺🇸
    Kinh doanh, Quản lý & Tài chính#4Phần mềm & Dịch vụ CNTT#8Khoa học Sự sống, Vật lý & Xã hội#9Pháp lý & Chính phủ#9

    Muse Spark 1.3 is Meta's proprietary coding and agentic model, released September 2, 2026 with a 1M-token context window and text, image, and video input, served through Muse Code and the Meta Model API. It is Meta's largest coding and agentic improvement to date, running roughly 20% fewer tool calls and 25% fewer tokens per task than Muse Spark 1.2 while sustaining longer-horizon work, and defaults to its broadly available xhigh reasoning level.

    Sep 2, 2026·ngữ cảnh 1M·mức chi phí 0.9x
  • Mô hình

    Gemini 3.8 Flash

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#9Khoa học Sự sống, Vật lý & Xã hội#10Giải trí, Thể thao & Truyền thông#10Y học & Chăm sóc sức khỏe#10

    Gemini 3.8 Flash is Google's workhorse Flash model, released September 2, 2026 and built on Gemini 3.7 Flash, with a 1M-token context window, 64K max output, tunable thinking levels, and multimodal input across text, image, audio, video, and PDF. It keeps the same context window and pricing as its predecessor while notably improving coding and reasoning and reducing hallucinations, making it a strong low-latency choice for high-volume agentic and everyday tasks.

    Sep 2, 2026·ngữ cảnh 1M·mức chi phí 0.8x
  • Mô hình

    Gemini 3.8 Flash (batch)

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#9Khoa học Sự sống, Vật lý & Xã hội#10Giải trí, Thể thao & Truyền thông#10Y học & Chăm sóc sức khỏe#10

    Gemini 3.8 Flash is Google's workhorse Flash model built on Gemini 3.7 Flash, with a 1M-token context window, tunable thinking levels, and multimodal input, tuned for low-latency agentic and everyday tasks. This is the batch variant for asynchronous, lower-cost processing.

    Sep 2, 2026·ngữ cảnh 1M·mức chi phí 0.4x
  • Mô hình

    Claude Fable 5.1

    Anthropic🇺🇸
    Phân tích tài liệu#2Y học & Chăm sóc sức khỏe#3Toán học#4Khoa học Sự sống, Vật lý & Xã hội#6

    Claude Fable 5.1 is Anthropic's most capable model, released September 2026 as an upgrade to Claude Fable 5 for coding, knowledge work, and long-running problem-solving, with a 1M-token context window. It improves agentic coding by over 30% and doubles its predecessor's Terminal-Bench-Science score while cutting costs through cheaper cache reads, and is the openly available counterpart to the more tightly safeguarded Mythos 5.1.

    Sep 1, 2026·ngữ cảnh 1M·mức chi phí 10x
  • Mô hình

    Claude Fable 5.1 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#2Y học & Chăm sóc sức khỏe#3Toán học#4Khoa học Sự sống, Vật lý & Xã hội#6

    Batch-processing variant of Claude Fable 5.1, Anthropic's most capable model for coding, knowledge work, and long-running problem-solving with a 1M-token context window. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Sep 1, 2026·ngữ cảnh 1M·mức chi phí 5x
  • Mô hình

    Granite 4.2 8B

    IBM🇺🇸
    Phần mềm & Dịch vụ CNTT#262Kinh doanh, Quản lý & Tài chính#281Y học & Chăm sóc sức khỏe#284Khoa học Sự sống, Vật lý & Xã hội#297

    Granite 4.2 8B is IBM's dense, decoder-only 8B-parameter reasoning model (Apache 2.0 license, 128K native context extendable to 512K) with switchable full, low-effort, and non-thinking modes, released August 25, 2026 as the mid-size member of the Granite 4.2 family. IBM positions it for enterprise use in mathematics, code generation, multilingual dialogue across 12 languages, and agentic tool-calling workflows on modest hardware.

    Aug 31, 2026·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Hy4 preview

    Tencent🇨🇳

    Hy4 Preview is Tencent Hunyuan's open-weight mixture-of-experts flagship (770B total parameters, 49B active, Apache 2.0 license, 1M-token context), released August 28, 2026 as a preview ahead of a later official release. It targets coding agents, complex tool-use workflows, and productivity tasks like office work, game development, and scientific research rather than general chat.

    Aug 28, 2026·ngữ cảnh 1M·mức chi phí 0.5x
  • Mô hình

    Ling 3.0 Flash Fin

    InclusionAI🇨🇳

    Ling-3.0-flash-fin is inclusionAI's (Ant Group) finance-tuned variant of Ling-3.0-flash, an open-weight 124B-parameter mixture-of-experts model (about 5.1B active per token, 256K-token context). It is optimized for financial research and multi-step investment workflows with heavy tool use, while retaining the base model's general reasoning, coding, and math abilities.

    Aug 27, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GLM Flash Latest

    Z.AI🇨🇳
    Toán học#25Phần mềm & Dịch vụ CNTT#26Hiểu hình ảnh#30Khoa học Sự sống, Vật lý & Xã hội#34

    This is an alias that always points to Z.ai's latest GLM Flash model.

    Aug 27, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Qwen3.8 Flash

    Qwen🇨🇳

    Qwen3.8 Flash is Alibaba's speed- and cost-optimized model in the Qwen3.8 family (1M-token context, up to 131K output tokens) with native text, image, and video input plus tool calling and structured outputs, released August 26, 2026. It sits below the flagship Qwen3.8 Max and is aimed at coding assistance, agentic workflows, document and codebase analysis, and long-video understanding where speed and cost matter more than peak capability.

    Aug 26, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    GLM 5.3 Flash

    Z.AI🇨🇳
    Toán học#25Phần mềm & Dịch vụ CNTT#26Hiểu hình ảnh#30Khoa học Sự sống, Vật lý & Xã hội#34

    GLM-5.3-Flash is Z.ai's first natively multimodal model in the GLM-5 family (320B total parameters, 18B active mixture-of-experts, MIT license, 1M-token context), released August 26, 2026 with a hybrid sparse-plus-linear attention architecture that keeps long-context behavior accurate while cutting compute overhead. It is built for efficient coding and long-horizon agent tasks, outperforming GLM-5.2 at roughly one-tenth the price while approaching flagship-class coding and agentic benchmark performance.

    Aug 26, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    GLM 5.3 Flash (batch)

    Z.AI🇨🇳
    Toán học#25Phần mềm & Dịch vụ CNTT#26Hiểu hình ảnh#30Khoa học Sự sống, Vật lý & Xã hội#34

    This is the batch-processing variant of GLM-5.3-Flash, Z.ai's open-weight natively multimodal mixture-of-experts model (320B total parameters, 18B active, MIT license, 1M-token context, hybrid sparse-plus-linear attention), offering the same coding and agentic capabilities at a lower cost via OpenRouter's asynchronous batch endpoint. It is best for high-volume, non-latency-sensitive coding or agent workloads where cost efficiency matters more than immediate response time.

    Aug 26, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Muse Spark 1.2 Contributor

    Meta🇺🇸

    Muse Spark 1.2 Contributor is Meta Superintelligence Labs' cost-reduced access tier for its proprietary Muse Spark 1.2 reasoning model (a Stable LatentMoE design activating 16 of 896 experts per token via Kimi Delta Attention), offering a roughly 1M-token context window across text, image, video, audio, and PDF input. In exchange for pricing more than 10 times cheaper than the standard tier, prompts and outputs on this tier may be used by Meta to improve its products, making it best suited for experimentation and early-stage projects rather than data-sensitive production use.

    Aug 21, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    DeepSeek V4 Flash Vision Exp

    DeepSeek🇨🇳

    DeepSeek V4 Flash Vision Exp is an experimental multimodal variant of DeepSeek V4 Flash 0731, a sparse MoE model (284B total, 13B active parameters) with a 1M-token context window that now accepts image input alongside text. It matches V4 Flash on agentic reasoning and world knowledge while adding document and chart understanding, visual question answering, and multimodal agent workflows, positioning it as the multimodal entry in DeepSeek's lightweight Flash line.

    Aug 21, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Hy-MT2-1.8B

    Tencent🇨🇳

    Hy-MT2-1.8B is Tencent Hunyuan's smallest open-weight translation model (1.8B dense parameters, Apache 2.0 license), supporting 33 languages as the lightest of the three Hy-MT2 'fast-thinking' sizes alongside the 7B dense and 30B-A3B MoE variants. Despite its size it beats mainstream commercial translation APIs on general and business-domain benchmarks, and AngelSlim quantization down to about 440MB makes it deployable on mobile chipsets for on-device translation.

    Aug 20, 2026·ngữ cảnh 8k·mức chi phí 0.1x
  • Mô hình

    Hy-MT2-30B-A3B

    Tencent🇨🇳

    Hy-MT2-30B-A3B is Tencent Hunyuan's largest translation model, a sparse mixture-of-experts architecture (30B total parameters, 3B active, 128 routed plus 1 shared expert) supporting 33 languages, released open-weight (Apache 2.0) alongside the smaller 1.8B and 7B dense Hy-MT2 variants. In fast-thinking mode it outperforms larger open-source general models such as DeepSeek-V4-Pro and Kimi K2.6 on translation benchmarks, making it the pick for high-quality business and domain-specific translation where the dense variants fall short.

    Aug 20, 2026·ngữ cảnh 8k·mức chi phí 0.1x
  • Mô hình

    GLM Latest

    Z.AI🇨🇳
    Toán học#18Khoa học Sự sống, Vật lý & Xã hội#19Y học & Chăm sóc sức khỏe#25Phần mềm & Dịch vụ CNTT#29

    This is an alias that always points to Z.ai's latest GLM model.

    Aug 19, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Hy-MT2-7B

    Tencent🇨🇳

    Hy-MT2-7B is Tencent's 7B-parameter translation model (8,192-token context), released August 19, 2026 as the mid-size entry in the Hy-MT2 fast-thinking translation family alongside 1.8B and 30B-A3B MoE siblings. It supports 33 language pairs plus Chinese dialect and minority-language pairs, with structured, delimiter-based, contextual, glossary-based, and style-guided translation workflows for dedicated multilingual translation rather than general-purpose chat.

    Aug 19, 2026·ngữ cảnh 8k·mức chi phí 0.1x
  • Mô hình

    GLM 5.3

    Z.AI🇨🇳
    Toán học#18Khoa học Sự sống, Vật lý & Xã hội#19Y học & Chăm sóc sức khỏe#25Phần mềm & Dịch vụ CNTT#29

    GLM-5.3 is Z.ai's open-weight flagship model (743B parameters, 1M-token context), an incremental upgrade over GLM-5.2 with stronger coding and much better token efficiency per task. Reasoning is always on with low, high, and max effort levels (max by default), and Z.ai positions it as a top open-weight model for complex software engineering and long-horizon agent tasks.

    Aug 18, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    GLM 5.3 (batch)

    Z.AI🇨🇳
    Toán học#18Khoa học Sự sống, Vật lý & Xã hội#19Y học & Chăm sóc sức khỏe#25Phần mềm & Dịch vụ CNTT#29

    This is the batch-processing variant of Z.ai's open-weight GLM-5.3 flagship model, offering the same 743B-parameter, 1M-token coding and long-horizon agentic capabilities at lower cost for asynchronous workloads.

    Aug 18, 2026·ngữ cảnh 1M·mức chi phí 0.4x
  • Mô hình

    Qwen3.8 27B

    Qwen🇨🇳
    Hiểu hình ảnh#58Phần mềm & Dịch vụ CNTT#76Toán học#80Khoa học Sự sống, Vật lý & Xã hội#83

    Qwen3.8 27B is a smaller, Apache 2.0-licensed open-weight model from Alibaba's Qwen3.8 generation, suited for general-purpose chat, coding, and agentic tasks at moderate cost.

    Aug 14, 2026·ngữ cảnh 1M·mức chi phí 0.5x
  • Dots3-Note Preview (free)

    dots-studio

    Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio (280B total, 16B active) with a 512K token context window and multimodal input support. It is the lightest model in the dots3 family, suited for coding, long-context document processing, multimodal understanding, and multi-step agentic workflows; this is the free-tier variant.

    Aug 14, 2026·ngữ cảnh 512k·Miễn phí
  • Mô hình

    Gemini 3.7 Flash

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#6Giải trí, Thể thao & Truyền thông#6Hiểu hình ảnh#6Khoa học Sự sống, Vật lý & Xã hội#14

    Gemini 3.7 Flash is Google's most intelligent workhorse model yet for coding and agents, an iterative refinement of 3.6 Flash with notably higher accuracy on debugging and issue resolution at roughly half the cost. It is designed for responsive multi-step agentic workflows, coding, and everyday knowledge work.

    Aug 13, 2026·ngữ cảnh 1M·mức chi phí 0.8x
  • Mô hình

    Gemini 3.7 Flash (batch)

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#6Giải trí, Thể thao & Truyền thông#6Hiểu hình ảnh#6Khoa học Sự sống, Vật lý & Xã hội#14

    Gemini 3.7 Flash is Google's most intelligent workhorse model yet for coding and agents, an iterative refinement of 3.6 Flash with notably higher accuracy on debugging and issue resolution. This is the batch variant for asynchronous, lower-cost processing.

    Aug 13, 2026·ngữ cảnh 1M·mức chi phí 0.4x
  • Mô hình

    Seed 2.1 Turbo

    Bytedance🇨🇳

    Seed 2.1 Turbo is a multimodal model from ByteDance Seed built for coding and long-horizon agent workflows, including end-to-end software delivery, multi-step task execution, and video/image understanding. It supports planning, debugging, and self-correction with a 262K-token context window.

    Aug 12, 2026·ngữ cảnh 262k·mức chi phí 0.5x
  • Mô hình

    Qwen3.8 2.4T A95B

    Qwen🇨🇳

    Qwen3.8 2.4T-A95B is Alibaba's massive open-weight mixture-of-experts flagship (2.4 trillion total parameters, 95B active), representing the top of the Qwen3.8 generation for frontier-level reasoning, coding, and agentic tasks.

    Aug 12, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    Seed-2.0-Code

    Bytedance🇨🇳

    Seed 2.0 Code is ByteDance Seed's model optimized specifically for agentic coding, suited to frontend development and multilingual programming tasks. It is designed to work well inside coding-agent tools such as Claude Code, Kilo, and OpenCode.

    Aug 12, 2026·ngữ cảnh 262k·mức chi phí 0.6x
  • Mô hình

    DeepSeek V4 Pro 0813

    DeepSeek🇨🇳
    Pháp lý & Chính phủ#37Kinh doanh, Quản lý & Tài chính#44Viết, Văn học & Ngôn ngữ#47Khoa học Sự sống, Vật lý & Xã hội#48

    DeepSeek V4 Pro (0813) is the August 13, 2026 general-availability checkpoint of DeepSeek V4 Pro, adding DSpark speculative decoding, reasoning-effort levels, and native OpenAI Responses API support. It is DeepSeek's flagship model for coding, reasoning, and agentic tasks in the V4 generation.

    Aug 12, 2026·ngữ cảnh 1M·mức chi phí 0.4x
  • Mô hình

    Grok 4.6

    xAI🇺🇸
    Phân tích tài liệu#25Toán học#41Hiểu hình ảnh#41Giải trí, Thể thao & Truyền thông#47

    Grok 4.6 is xAI's flagship model built for long-running agentic work and ambitious interactive and visual tasks, with a 500,000-token context window and text/image input. It is tuned for advanced reasoning, coding, and complex agent workflows at the top of xAI's model lineup.

    Aug 12, 2026·ngữ cảnh 500k·mức chi phí 1x
  • Mô hình

    LFM2.5-2.6B (free)

    Liquid🇺🇸

    LFM2.5-2.6B is Liquid AI's on-device agentic model (2.6B parameters, 128K context), pre-trained on ~34T tokens to plan, call tools, and run multi-step tasks entirely on phones, laptops, and edge devices with no data leaving the device. It leads its size class on instruction-following and tool-use benchmarks while running in under 2.5 GB of memory; this is OpenRouter's rate-limited free endpoint.

    Aug 11, 2026·ngữ cảnh 66k·Miễn phí
  • Mô hình

    Nemotron 3.5 Lightning

    NVIDIA🇺🇸

    Nemotron 3.5 Lightning is NVIDIA's open 30B mixture-of-experts model with 3B active parameters and a 1M-token context window, built for fast, accurate execution within long-running agent pipelines. It is designed as a low-latency, low-cost execution layer for always-on agentic systems.

    Aug 11, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Nemotron 3.5 Lightning (free)

    NVIDIA🇺🇸

    Nemotron 3.5 Lightning is NVIDIA's open 30B mixture-of-experts model with 3B active parameters and a 1M-token context window, built as a low-latency, low-cost execution layer for always-on agentic systems. This is the free-tier variant of the same model.

    Aug 11, 2026·ngữ cảnh 1M·Miễn phí
  • Mô hình

    Sakana Namazu

    Sakana AI🇯🇵

    Sakana Namazu is an OpenAI-compatible model from Sakana AI, built as a fine-tune of Moonshot AI's open-weight Kimi model and tailored for Japanese language and business use cases. It suits Japanese-language chat, business writing, and localized assistant tasks.

    Aug 11, 2026·ngữ cảnh 262k·mức chi phí 0.8x
  • Mô hình

    Solar Pro 4

    Upstage🇰🇷
    Toán học#133Phần mềm & Dịch vụ CNTT#144Kinh doanh, Quản lý & Tài chính#174Khoa học Sự sống, Vật lý & Xã hội#178

    Solar Pro 4 is Upstage's next-generation flagship model, an agent-first successor to Solar Pro 3 focused on further reasoning gains at low cost. It targets multi-step tool use and production-scale agentic workloads while retaining Upstage's efficiency-focused, compact-parameter design philosophy.

    Aug 10, 2026·ngữ cảnh 524k·mức chi phí 0.1x
  • Mô hình

    Muse Glimmer 30B

    Meta🇺🇸

    Muse Glimmer 30B is Meta's open-weight multimodal model, a distilled version of the larger Muse Spark model, combining an LLM with a dedicated perception encoder for text and visual input. It is built for local agentic and coding workloads, and can run entirely offline on a single consumer GPU.

    Aug 9, 2026·ngữ cảnh 131k·mức chi phí 0.3x
  • Mô hình

    Muse Spark 1.2

    Meta🇺🇸
    Pháp lý & Chính phủ#1Kinh doanh, Quản lý & Tài chính#2Y học & Chăm sóc sức khỏe#9Hiểu hình ảnh#9

    Muse Spark 1.2 is an updated release in Meta Superintelligence Labs' Muse family, sharing Muse Spark's focus on multimodal reasoning, coding, and AI-assisted software development with long-context support. It offers incremental improvements over Muse Spark 1.1 for agentic and coding workflows.

    Aug 5, 2026·ngữ cảnh 1M·mức chi phí 0.9x
  • Mô hình

    DeepSeek V4 Flash Latest

    DeepSeek🇨🇳
    Y học & Chăm sóc sức khỏe#93Pháp lý & Chính phủ#94Khoa học Sự sống, Vật lý & Xã hội#95Viết, Văn học & Ngôn ngữ#99

    This is an alias that always points to DeepSeek's latest V4-Flash model.

    Aug 1, 2026·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    DeepSeek V4 Flash 0731

    DeepSeek🇨🇳
    Y học & Chăm sóc sức khỏe#93Pháp lý & Chính phủ#94Khoa học Sự sống, Vật lý & Xã hội#95Viết, Văn học & Ngôn ngữ#99

    DeepSeek V4 Flash (0731) is the July 31, 2026 public-beta checkpoint of DeepSeek's V4 Flash model, DeepSeek's fast, cost-efficient V4-generation model tuned for agentic workflows. It offers the same lightweight, high-throughput design as other V4 Flash releases.

    Jul 31, 2026·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    Inkling Small

    Thinking Machines Lab🇺🇸
    Hiểu hình ảnh#81Toán học#120Phần mềm & Dịch vụ CNTT#121Kinh doanh, Quản lý & Tài chính#143

    Inkling-Small is a compact open-weights sibling of Thinking Machines Lab's Inkling model, at 276B total parameters (12B active) versus Inkling's 975B. It reasons natively over text, images, and audio with adjustable thinking effort, achieving performance close to the full Inkling model at a fraction of its size.

    Jul 30, 2026·ngữ cảnh 524k·mức chi phí 0.3x
  • Mô hình

    Inkling Small (free)

    Thinking Machines Lab🇺🇸
    Hiểu hình ảnh#81Toán học#120Phần mềm & Dịch vụ CNTT#121Kinh doanh, Quản lý & Tài chính#143

    This is the free OpenRouter endpoint for Thinking Machines Lab's Inkling-Small, a 276B-parameter mixture-of-experts model (12B active, Apache 2.0 license, 1M-token context) released July 30, 2026 as the smaller, efficient sibling in the Inkling family. Free access is rate-limited to agentic-harness use with prompt and output logging, making it best suited for evaluating its coding, agentic, and RAG capabilities rather than production workloads.

    Jul 30, 2026·ngữ cảnh 1M·Miễn phí
  • Mô hình

    Qwen3.7 Flash

    Qwen🇨🇳

    Qwen3.7 Flash is Alibaba's fast, low-cost proprietary tier of the API-only Qwen3.7 generation, built for high-throughput, latency-sensitive tasks.

    Jul 27, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Claude Opus 5

    Anthropic🇺🇸
    Phân tích tài liệu#1Toán học#2Pháp lý & Chính phủ#10Y học & Chăm sóc sức khỏe#13

    Claude Opus 5 is Anthropic's flagship model in the Claude 5 generation, representing the top of Anthropic's intelligence tier for coding, agentic workflows, and complex reasoning. It targets the most demanding professional software engineering and long-horizon agent tasks Anthropic offers.

    Jul 24, 2026·ngữ cảnh 1M·mức chi phí 5x
  • Mô hình

    Claude Opus 5 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#1Toán học#2Pháp lý & Chính phủ#10Y học & Chăm sóc sức khỏe#13

    Batch-processing variant of Claude Opus 5, Anthropic's Claude 5 generation flagship model. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Jul 24, 2026·ngữ cảnh 1M·mức chi phí 2.5x
  • Mô hình

    Ling 3.0 Flash

    InclusionAI🇨🇳

    Ling 3.0 Flash is InclusionAI's (Ant Group) successor to Ling 2.6 Flash, a hybrid-reasoning Mixture-of-Experts model (~124B total, ~5.1B active) combining Kimi Delta Attention with Multi-Head Latent Attention for efficient long-range memory. It supports both thinking and non-thinking modes with a native 262K context, suited for fast general-purpose coding and agentic tasks.

    Jul 23, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Laguna S 2.1

    Poolside🇺🇸

    Laguna S 2.1 is an open-weight agentic coding model from Poolside, a 118B-parameter mixture-of-experts system with 8B active parameters per token and up to a 1M-token context window. Trained with reinforcement learning inside Poolside's own agent harness, it targets agentic software engineering tasks and is competitive with much larger models on SWE-bench and Terminal-Bench.

    Jul 21, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Laguna S 2.1 (free)

    Poolside🇺🇸

    Free-tier variant of Poolside's Laguna S 2.1, an open-weight 118B MoE agentic coding model with a 1M-token context window trained via RL in Poolside's agent harness. Same underlying model as the paid version, suited for agentic software engineering and terminal/coding tasks.

    Jul 21, 2026·ngữ cảnh 262k·Miễn phí
  • Mô hình

    Gemini 3.6 Flash

    Google🇺🇸
    Giải trí, Thể thao & Truyền thông#17Viết, Văn học & Ngôn ngữ#19Phân tích tài liệu#24Hiểu hình ảnh#25

    Gemini 3.6 Flash is Google's workhorse model delivering a step up in coding, knowledge work, and multimodal performance over 3.5 Flash, with meaningfully improved token efficiency. It targets everyday agentic and coding tasks that need strong quality at Flash-tier cost.

    Jul 21, 2026·ngữ cảnh 1M·mức chi phí 0.8x
  • Mô hình

    Gemini 3.6 Flash (batch)

    Google🇺🇸
    Giải trí, Thể thao & Truyền thông#17Viết, Văn học & Ngôn ngữ#19Phân tích tài liệu#24Hiểu hình ảnh#25

    Gemini 3.6 Flash is Google's workhorse model delivering a step up in coding, knowledge work, and multimodal performance with improved token efficiency. This is the batch variant for asynchronous, lower-cost processing.

    Jul 21, 2026·ngữ cảnh 1M·mức chi phí 0.4x
  • Mô hình

    Gemini 3.5 Flash Lite

    Google🇺🇸
    Hiểu hình ảnh#40Khoa học Sự sống, Vật lý & Xã hội#61Giải trí, Thể thao & Truyền thông#64Viết, Văn học & Ngôn ngữ#68

    Gemini 3.5 Flash-Lite is the most cost-effective model in Google's 3.5 Flash line, built for high-volume, low-latency tasks where price and speed are the priority. It trades peak reasoning depth for efficiency at scale.

    Jul 21, 2026·ngữ cảnh 1M·mức chi phí 0.5x
  • Mô hình

    Gemini 3.5 Flash Lite (batch)

    Google🇺🇸
    Hiểu hình ảnh#40Khoa học Sự sống, Vật lý & Xã hội#61Giải trí, Thể thao & Truyền thông#64Viết, Văn học & Ngôn ngữ#68

    Gemini 3.5 Flash-Lite is the most cost-effective model in Google's 3.5 Flash line, built for high-volume, low-latency tasks where price and speed are the priority. This is the batch variant for asynchronous, lower-cost processing.

    Jul 21, 2026·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    LongCat 2.0

    Meituan🇨🇳

    LongCat-2.0 is Meituan's open-weight 1.6-trillion-parameter Mixture-of-Experts model (~48B active) with a native 1M-token context window, using sparse attention for efficient long-context processing. It targets agentic coding, repository-level code changes, and long-horizon problem solving, with reported performance comparable to Gemini 3.1 Pro on coding benchmarks.

    Jul 20, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    Inkling

    Thinking Machines Lab🇺🇸
    Toán học#63Phần mềm & Dịch vụ CNTT#90Khoa học Sự sống, Vật lý & Xã hội#92Kinh doanh, Quản lý & Tài chính#92

    Inkling is Thinking Machines Lab's first flagship model, a 975B-parameter (41B active) open-weights mixture-of-experts model that reasons natively over text, images, and audio with up to a 1M-token context window. It is designed as a general-purpose frontier model for multimodal reasoning and agentic tasks, released under Apache 2.0.

    Jul 17, 2026·ngữ cảnh 524k·mức chi phí 0.8x
  • Mô hình

    Inkling (free)

    Thinking Machines Lab🇺🇸
    Toán học#63Phần mềm & Dịch vụ CNTT#90Khoa học Sự sống, Vật lý & Xã hội#92Kinh doanh, Quản lý & Tài chính#92

    This is the free OpenRouter endpoint for Thinking Machines Lab's Inkling, a 975B-parameter mixture-of-experts model (41B active, Apache 2.0 license, 1M-token context) released July 2026 as the flagship of the Inkling family and the company's first in-house model. Free access is rate-limited to agentic-harness use with prompt and output logging, making it best suited for evaluating general-purpose reasoning, coding, agentic tool-use, and RAG rather than production workloads.

    Jul 17, 2026·ngữ cảnh 1M·Miễn phí
  • Mô hình

    Auto Router (Beta)

    OpenRouter🇺🇸

    Auto Router (Beta) is the early-access track of OpenRouter's task-aware Auto Router, where new routing behaviors land before they reach the stable openrouter/auto. It uses the same market-driven selection (prompts classified into ~30 task types, routed by trailing 7-day community spend with a cost_tier dial) and prices each response at the routed model's rate; choose it to try the newest routing changes first.

    Jul 17, 2026·ngữ cảnh 2M·mức chi phí -
  • Mô hình

    Kimi K3

    Moonshot🇨🇳
    Pháp lý & Chính phủ#3Khoa học Sự sống, Vật lý & Xã hội#5Y học & Chăm sóc sức khỏe#8Toán học#12

    Kimi K3 is Moonshot AI's July 2026 flagship model, a 2.8-trillion-parameter mixture-of-experts model with native multimodal vision and a 1M-token context window. It is built for long-horizon coding, reasoning, and agent workflows, positioned as an open-weight competitor to top proprietary frontier models.

    Jul 16, 2026·ngữ cảnh 1M·mức chi phí 2.5x
  • Mô hình

    Kimi K3 (batch)

    Moonshot🇨🇳
    Pháp lý & Chính phủ#3Khoa học Sự sống, Vật lý & Xã hội#5Y học & Chăm sóc sức khỏe#8Toán học#12

    This is the asynchronous batch-processing endpoint for Moonshot AI's Kimi K3, a 2.8T-parameter open-weight mixture-of-experts reasoning model (1M-token context) released July 2026 under a custom revenue-tiered Kimi K3 License. Unlike typical batch endpoints this one carries no price discount over the standard endpoint, so it is worth reaching for only when queued, asynchronous processing itself is the requirement rather than for cost savings.

    Jul 16, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Muse Spark 1.1

    Meta🇺🇸
    Kinh doanh, Quản lý & Tài chính#9Phần mềm & Dịch vụ CNTT#14Pháp lý & Chính phủ#16Toán học#17

    Muse Spark 1.1 is a large language model from Meta Superintelligence Labs, the first model in Meta's Muse family, designed for multimodal reasoning, coding, and AI-assisted software development with up to a million tokens of context. It targets complex agentic and coding workflows requiring long-context understanding.

    Jul 16, 2026·ngữ cảnh 1M·mức chi phí 0.9x
  • Mô hình

    KAT-Coder-Pro V2.5

    Kwaipilot🇨🇳

    KAT-Coder-Pro V2.5 is Kwaipilot's flagship agentic coding model, trained on 100,000+ verifiable repository environments and capable of autonomously handling an entire issue or business workflow end to end in a real codebase. It ranks highly on agentic coding benchmarks like PinchBench and SWE-Bench Pro.

    Jul 10, 2026·ngữ cảnh 262k·mức chi phí 0.6x
  • Mô hình

    GPT-5.6 Luna Pro

    OpenAI🇺🇸

    GPT-5.6 Luna Pro is a higher-reasoning-effort version of the Luna tier in OpenAI's GPT-5.6 family, aimed at latency-sensitive chat, classification, and lightweight agentic tasks that need a bit more accuracy than the base Luna model.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 0.2x
  • Mô hình

    GPT-5.6 Luna Pro (batch)

    OpenAI🇺🇸

    GPT-5.6 Luna Pro is a higher-reasoning-effort version of the fast, low-cost Luna tier in OpenAI's GPT-5.6 family. This is the batch variant for cheaper asynchronous processing.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    GPT-5.6 Luna

    OpenAI🇺🇸
    Phân tích tài liệu#23Toán học#39Hiểu hình ảnh#44Kinh doanh, Quản lý & Tài chính#54

    GPT-5.6 Luna is the fast, low-cost tier of OpenAI's GPT-5.6 family, tuned for high-volume, latency-sensitive work like chat, classification, and lightweight agents while retaining genuine GPT-5.6 reasoning.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 0.2x
  • Mô hình

    GPT-5.6 Luna (batch)

    OpenAI🇺🇸
    Phân tích tài liệu#23Toán học#39Hiểu hình ảnh#44Kinh doanh, Quản lý & Tài chính#54

    GPT-5.6 Luna is the fast, low-cost tier of OpenAI's GPT-5.6 family for high-volume chat, classification, and lightweight agents. This is the batch variant for cheaper asynchronous processing.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    GPT-5.6 Terra Pro

    OpenAI🇺🇸

    GPT-5.6 Terra Pro is a higher-reasoning-effort version of the balanced Terra tier in OpenAI's GPT-5.6 family, for everyday coding, reasoning, and agentic tasks that need a bit more accuracy.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT-5.6 Terra Pro (batch)

    OpenAI🇺🇸

    GPT-5.6 Terra Pro is a higher-reasoning-effort version of the balanced Terra tier in OpenAI's GPT-5.6 family. This is the batch variant for cheaper asynchronous processing.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 1x
  • Mô hình

    GPT-5.6 Terra

    OpenAI🇺🇸
    Phân tích tài liệu#13Toán học#33Hiểu hình ảnh#36Pháp lý & Chính phủ#43

    GPT-5.6 Terra is the balanced, everyday tier of OpenAI's GPT-5.6 family, offering coding, reasoning, and agentic performance competitive with GPT-5.5 at roughly half the cost.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT-5.6 Terra (batch)

    OpenAI🇺🇸
    Phân tích tài liệu#13Toán học#33Hiểu hình ảnh#36Pháp lý & Chính phủ#43

    GPT-5.6 Terra is the balanced, everyday tier of OpenAI's GPT-5.6 family, offering coding, reasoning, and agentic performance at low cost. This is the batch variant for cheaper asynchronous processing.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 1x
  • Mô hình

    GPT-5.6 Sol Pro

    OpenAI🇺🇸

    GPT-5.6 Sol Pro is an extended-reasoning version of the Sol flagship tier in OpenAI's GPT-5.6 family, aimed at the hardest complex reasoning, multi-step coding, and long-horizon agentic tasks.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT-5.6 Sol Pro (batch)

    OpenAI🇺🇸

    GPT-5.6 Sol Pro is an extended-reasoning version of the Sol flagship tier in OpenAI's GPT-5.6 family, for the hardest reasoning and agentic coding tasks. This is the batch variant for cheaper asynchronous processing.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 1x
  • Mô hình

    GPT-5.6 Sol

    OpenAI🇺🇸
    Toán học#10Phân tích tài liệu#10Viết, Văn học & Ngôn ngữ#11Giải trí, Thể thao & Truyền thông#11

    GPT-5.6 Sol is the flagship tier of OpenAI's GPT-5.6 family, built for complex reasoning, multi-step coding, and long-horizon agentic work, with particular strength on command-line and terminal-driven tasks.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 2x
  • Mô hình

    GPT-5.6 Sol (batch)

    OpenAI🇺🇸
    Toán học#10Phân tích tài liệu#10Viết, Văn học & Ngôn ngữ#11Giải trí, Thể thao & Truyền thông#11

    GPT-5.6 Sol is the flagship tier of OpenAI's GPT-5.6 family, built for complex reasoning, multi-step coding, and long-horizon agentic work. This is the batch variant for cheaper asynchronous processing.

    Jul 9, 2026·ngữ cảnh 1.1M·mức chi phí 1x
  • Mô hình

    Grok 4.5

    xAI🇺🇸
    Phân tích tài liệu#26Toán học#28Hiểu hình ảnh#28Pháp lý & Chính phủ#41

    Grok 4.5 is a flagship xAI model positioned for code generation and general high-intelligence use cases, sitting between Grok 4.3 and Grok 4.6 in xAI's rapid release cadence. It is designed for demanding reasoning, coding, and agentic workloads.

    Jul 8, 2026·ngữ cảnh 500k·mức chi phí 1x
  • Mô hình

    Grok Latest

    xAI🇺🇸
    Pháp lý & Chính phủ#72Y học & Chăm sóc sức khỏe#72Viết, Văn học & Ngôn ngữ#80Giải trí, Thể thao & Truyền thông#85

    This is an alias that always points to xAI's latest flagship Grok model.

    Jul 8, 2026·ngữ cảnh 500k·mức chi phí 1x
  • Mô hình

    Aion-3.0-Mini

    Aion Labs🇵🇱

    Aion-3.0-mini is a smaller, faster member of AionLabs' Aion-3.0 roleplay and storytelling family, built on DeepSeek models with a similar collaborative multi-model generation approach. It targets the same creative-writing and roleplay use cases at lower cost and latency.

    Jul 7, 2026·ngữ cảnh 131k·mức chi phí 0.3x
  • Mô hình

    Aion-3.0

    Aion Labs🇵🇱

    Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs built on the GLM family, using a collaborative generation process where multiple specialized models contribute to each response. It is designed for rich narrative structure and compelling character-driven fiction.

    Jul 7, 2026·ngữ cảnh 131k·mức chi phí 1.5x
  • Mô hình

    Hy3

    Tencent🇨🇳
    Toán học#23Khoa học Sự sống, Vật lý & Xã hội#43Y học & Chăm sóc sức khỏe#49Phần mềm & Dịch vụ CNTT#74

    Hy3 is Tencent Hunyuan's open-source flagship model, a 295B-parameter MoE (about 21B active) with a hybrid fast/slow reasoning system and 256k context. It targets coding, research, and agentic product integrations, emphasizing reliable tool calling and long multi-turn conversations.

    Jul 6, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Laguna XS 2.1

    Poolside🇺🇸

    Laguna XS 2.1 is Poolside's smaller open-weight coding model (33B total parameters MoE), built for local, lightweight agentic coding workloads. It applies learnings from Poolside's larger Laguna models while staying fast and cheap enough to run on modest hardware.

    Jul 2, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Laguna XS 2.1 (free)

    Poolside🇺🇸

    Free-tier variant of Poolside's Laguna XS 2.1, a compact open-weight MoE coding model designed for local agentic coding. Same underlying model as the paid version, aimed at lightweight, cost-effective code generation and tool-use tasks.

    Jul 2, 2026·ngữ cảnh 262k·Miễn phí
  • Mô hình

    Claude Sonnet 5

    Anthropic🇺🇸
    Phân tích tài liệu#17Khoa học Sự sống, Vật lý & Xã hội#38Toán học#38Hiểu hình ảnh#39

    Claude Sonnet 5 is Anthropic's most capable Sonnet-tier model, bringing near-Opus intelligence to coding, agentic workflows, and professional knowledge work at Sonnet-level cost. It is a strict improvement over Sonnet 4.6 across the cost-performance curve and serves as the default model for Anthropic's Free and Pro plans.

    Jun 30, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Claude Sonnet 5 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#17Khoa học Sự sống, Vật lý & Xã hội#38Toán học#38Hiểu hình ảnh#39

    Batch-processing variant of Claude Sonnet 5, Anthropic's near-Opus-capable Sonnet-tier flagship. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Jun 30, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    Fugu Ultra

    Sakana AI🇯🇵

    Sakana Fugu Ultra is Sakana AI's multi-agent orchestration model that coordinates a pool of frontier language models through a single API endpoint, dynamically assigning roles via learned orchestration. It targets complex, multi-step tasks like agentic software engineering, achieving top-tier benchmark scores by intelligently routing work across underlying models.

    Jun 24, 2026·ngữ cảnh 1M·mức chi phí 6x
  • Mô hình

    North Mini Code (free)

    Cohere🇨🇦

    North Mini Code is Cohere's first developer-focused coding model, a 30B-total/3B-active mixture-of-experts model released open-weight under Apache 2.0. It is trained for agentic coding, terminal tasks, and tool use across agent harnesses like OpenCode and SWE-Agent, with a 256K context window; this is the free-tier variant of the same model.

    Jun 17, 2026·ngữ cảnh 256k·Miễn phí
  • Mô hình

    GLM 5.2

    Z.AI🇨🇳
    Khoa học Sự sống, Vật lý & Xã hội#24Pháp lý & Chính phủ#25Toán học#27Y học & Chăm sóc sức khỏe#30

    GLM-5.2 is Z.ai's flagship open-source model, shipping with a 1M-token context window, two selectable reasoning effort levels, and an unrestricted MIT license. It is competitive with top closed models on demanding software-engineering benchmarks, trailing Claude Opus-class models only narrowly while outperforming several rivals on coding tasks.

    Jun 16, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Kimi K2.7 Code

    Moonshot🇨🇳

    Kimi K2.7 Code is Moonshot AI's June 2026 specialized coding model, tuned for stronger instruction-following over long contexts and higher success on coding tasks. It cuts thinking-token usage by about 30% versus K2.6 while improving benchmark scores, making it well suited for agentic coding workflows.

    Jun 12, 2026·ngữ cảnh 262k·mức chi phí 0.7x
  • Mô hình

    Claude Fable Latest

    Anthropic🇺🇸
    Phân tích tài liệu#2Y học & Chăm sóc sức khỏe#3Toán học#4Khoa học Sự sống, Vật lý & Xã hội#6

    This is an alias that always points to Anthropic's latest Claude Fable model.

    Jun 9, 2026·ngữ cảnh 1M·mức chi phí 10x
  • Mô hình

    Claude Fable 5

    Anthropic🇺🇸
    Hiểu hình ảnh#1Viết, Văn học & Ngôn ngữ#3Khoa học Sự sống, Vật lý & Xã hội#3Giải trí, Thể thao & Truyền thông#3

    Claude Fable 5 is Anthropic's most capable coding-focused model, built for long-running, ambiguous, asynchronous work such as legacy migrations and gnarly production bugs that can span hours or days. It self-corrects through verification loops and is state of the art on nearly all tested coding benchmarks, with a 1M-token context window.

    Jun 9, 2026·ngữ cảnh 1M·mức chi phí 10x
  • Mô hình

    Claude Fable 5 (batch)

    Anthropic🇺🇸
    Hiểu hình ảnh#1Viết, Văn học & Ngôn ngữ#3Khoa học Sự sống, Vật lý & Xã hội#3Giải trí, Thể thao & Truyền thông#3

    Batch-processing variant of Claude Fable 5, Anthropic's most capable coding model for long-running, ambiguous, asynchronous engineering work. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Jun 9, 2026·ngữ cảnh 1M·mức chi phí 5x
  • Mô hình

    Nemotron 3.5 Content Safety

    NVIDIA🇺🇸

    Nemotron 3.5 Content Safety is a small NVIDIA moderation model built on Google's Gemma-3-4B and fine-tuned on multimodal, multilingual safety data. It classifies prompts and responses across 23 safety categories such as violence, hate, self-harm, and PII in 12 languages, and is meant for guardrailing other models rather than general chat.

    Jun 4, 2026·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Nemotron 3.5 Content Safety (free)

    NVIDIA🇺🇸

    Nemotron 3.5 Content Safety is a small NVIDIA moderation model built on Google's Gemma-3-4B and fine-tuned on multimodal, multilingual safety data. It classifies prompts and responses across 23 safety categories such as violence, hate, self-harm, and PII in 12 languages, and is meant for guardrailing other models rather than general chat.

    Jun 4, 2026·ngữ cảnh 128k·Miễn phí
  • Mô hình

    Nemotron 3 Ultra

    NVIDIA🇺🇸

    Nemotron 3 Ultra is NVIDIA's flagship open-weight model in the Nemotron 3 family, a 550B-parameter mixture-of-experts model with 55B active parameters. It is optimized for orchestrating complex, long-running agent workflows, combining frontier reasoning with high inference throughput.

    Jun 4, 2026·ngữ cảnh 262k·mức chi phí 0.5x
  • Mô hình

    Nemotron 3 Ultra (free)

    NVIDIA🇺🇸

    Nemotron 3 Ultra is NVIDIA's flagship open-weight model in the Nemotron 3 family, a 550B-parameter mixture-of-experts model with 55B active parameters optimized for complex, long-running agent workflows. This is the free-tier variant of the same model.

    Jun 4, 2026·ngữ cảnh 1M·Miễn phí
  • Mô hình

    Qwen3.7 Plus

    Qwen🇨🇳
    Phân tích tài liệu#30Hiểu hình ảnh#37Toán học#52Phần mềm & Dịch vụ CNTT#72

    Qwen3.7 Plus is Alibaba's multimodal agent-focused model in the API-only Qwen3.7 generation, combining vision and language understanding with strong agentic tool-use capability.

    Jun 3, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    MiniMax M3

    Minimax🇨🇳
    Phân tích tài liệu#32Hiểu hình ảnh#63Kinh doanh, Quản lý & Tài chính#78Phần mềm & Dịch vụ CNTT#85

    MiniMax-M3 is MiniMax's latest M-series model for agentic reasoning, tool use, coding, multimodal chat input, and long-context tasks. It represents the newest generation of MiniMax's flagship line, building on the coding and agentic strengths of the M2 series.

    May 31, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    Step 3.7 Flash

    Stepfun🇨🇳

    Step 3.7 Flash is StepFun's multimodal vision-language model, built on Step 3.5 Flash with an added vision encoder for native image understanding. With a 256k context window, selectable reasoning levels, and high throughput, it targets agentic, coding, and search workflows that mix text and visual input.

    May 28, 2026·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    Claude Opus 4.8

    Anthropic🇺🇸
    Phân tích tài liệu#12Kinh doanh, Quản lý & Tài chính#13Toán học#13Pháp lý & Chính phủ#15

    Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family at this point in the lineup, supporting text, image, and file input with reasoning and a 1M-token context window. It is aimed at the highest-difficulty coding, analysis, and agentic tasks.

    May 27, 2026·ngữ cảnh 1M·mức chi phí 5x
  • Mô hình

    Claude Opus 4.8 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#12Kinh doanh, Quản lý & Tài chính#13Toán học#13Pháp lý & Chính phủ#15

    Batch-processing variant of Claude Opus 4.8, Anthropic's most capable generally available Opus model. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    May 27, 2026·ngữ cảnh 1M·mức chi phí 2.5x
  • Mô hình

    Qwen3.7 Max

    Qwen🇨🇳

    Qwen3.7 Max is Alibaba's flagship API-only model (no open weights) in the Qwen3.7 generation, positioned as the text flagship for advanced reasoning and agentic tasks.

    May 21, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    Grok Build 0.1

    xAI🇺🇸

    Grok Build 0.1 is xAI's fast coding-specialist model, trained specifically for agentic software-engineering workflows like code generation, review, debugging, and multi-step development rather than general chat. It has a 256K context window with text and image input and is optimized for interactive coding agents and high-volume tool-use workloads.

    May 20, 2026·ngữ cảnh 256k·mức chi phí 0.5x
  • Mô hình

    Gemini 3.5 Flash

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#18Hiểu hình ảnh#19Phân tích tài liệu#20Giải trí, Thể thao & Truyền thông#24

    Gemini 3.5 Flash is Google's near-Pro intelligence model at Flash-tier cost and speed, delivering strong coding proficiency and parallel agentic execution for complex, long-horizon tasks. It is designed to give developers frontier-level agent and coding performance without Pro-level pricing.

    May 19, 2026·ngữ cảnh 1M·mức chi phí 1.5x
  • Mô hình

    Gemini 3.5 Flash (batch)

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#18Hiểu hình ảnh#19Phân tích tài liệu#20Giải trí, Thể thao & Truyền thông#24

    Gemini 3.5 Flash is Google's near-Pro intelligence model at Flash-tier cost and speed, delivering strong coding proficiency and parallel agentic execution for complex, long-horizon tasks. This is the batch variant for asynchronous, lower-cost processing.

    May 19, 2026·ngữ cảnh 1M·mức chi phí 0.9x
  • Mô hình

    Perceptron Mk1

    Perceptron🇺🇸

    Perceptron Mk1 is the flagship vision-language model from Perceptron AI, a startup founded by former Meta FAIR researchers, built for video understanding and embodied/spatial reasoning. It is designed for video QA, summarization and event detection, OCR and document parsing on messy real-world inputs, object detection/counting, and pose estimation, delivering frontier-competitive results at a much lower cost.

    May 12, 2026·ngữ cảnh 33k·mức chi phí 0.3x
  • Mô hình

    Gemini 3.1 Flash Lite

    Google🇺🇸

    Gemini 3.1 Flash-Lite is Google's high-efficiency multimodal model optimized for low-latency, high-volume workloads across text, image, video, audio, and PDF input. It is designed for lightweight agentic workflows and simple data extraction where responsiveness and cost matter most.

    May 7, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    Gemini 3.1 Flash Lite (batch)

    Google🇺🇸

    Gemini 3.1 Flash-Lite is Google's high-efficiency multimodal model optimized for low-latency, high-volume workloads across text, image, video, audio, and PDF input. This is the batch variant for asynchronous, lower-cost processing.

    May 7, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    GPT Chat Latest

    OpenAI🇺🇸

    GPT Chat Latest is an alias that always points to the current production ChatGPT model snapshot, tuned for natural conversational chat rather than heavy extended reasoning. It suits general assistant use where staying in sync with ChatGPT's live behavior matters.

    May 5, 2026·ngữ cảnh 400k·mức chi phí 6x
  • Mô hình

    Grok 4.3

    xAI🇺🇸
    Hiểu hình ảnh#56Giải trí, Thể thao & Truyền thông#86Viết, Văn học & Ngôn ngữ#87Phần mềm & Dịch vụ CNTT#93

    Grok 4.3 is an iteration in xAI's Grok 4.x line, improving on Grok 4.20 with incremental gains in reasoning, coding, and agentic task performance. It is positioned as a high-intelligence general-purpose model for complex problem solving.

    Apr 30, 2026·ngữ cảnh 1M·mức chi phí 0.6x
  • Mô hình

    Grok 4.3 (batch)

    xAI🇺🇸
    Hiểu hình ảnh#56Giải trí, Thể thao & Truyền thông#86Viết, Văn học & Ngôn ngữ#87Phần mềm & Dịch vụ CNTT#93

    This is the batch-processing variant of xAI's Grok 4.3, offering the same high-intelligence reasoning, coding, and agentic capabilities of the Grok 4.x line at lower cost for asynchronous workloads. Requests are processed on a delayed, non-interactive basis, making it best for large-scale jobs that do not need real-time responses.

    Apr 30, 2026·ngữ cảnh 1M·mức chi phí 0.5x
  • Mô hình

    Mistral Medium 3.5

    Mistral🇫🇷
    Hiểu hình ảnh#86Toán học#106Phần mềm & Dịch vụ CNTT#108Kinh doanh, Quản lý & Tài chính#113

    Mistral Medium 3.5 is Mistral AI's April 2026 flagship-tier model, a dense 128B model with a 256K context window that merges instruction-following, reasoning, and coding into a single set of weights. It is the recommended successor to Mistral Medium 3.1 for general-purpose and agentic tasks.

    Apr 30, 2026·ngữ cảnh 262k·mức chi phí 1.5x
  • Mô hình

    Mistral Medium 3.5 (batch)

    Mistral🇫🇷
    Hiểu hình ảnh#86Toán học#106Phần mềm & Dịch vụ CNTT#108Kinh doanh, Quản lý & Tài chính#113

    This is the batch-processing variant of Mistral Medium 3.5 (mistral-medium-3-5), Mistral AI's 128B-parameter dense model (Modified MIT license, 256K-token context) that merges the former Magistral reasoning and Devstral 2 coding models into one self-hostable system. It targets balanced, flagship-adjacent agentic and coding workloads with reliable multi-tool calling, and the batch endpoint trades real-time latency for lower cost on asynchronous, high-volume jobs.

    Apr 30, 2026·ngữ cảnh 262k·mức chi phí 0.8x
  • Mô hình

    Nemotron 3 Nano Omni (free)

    NVIDIA🇺🇸

    Nemotron 3 Nano Omni Reasoning is a multimodal, reasoning-tuned variant of NVIDIA's Nemotron 3 Nano, built for omni (text and image, likely audio) understanding combined with extended chain-of-thought. It targets efficient agentic applications that need both reasoning and multimodal input at low compute cost.

    Apr 28, 2026·ngữ cảnh 256k·Miễn phí
  • Mô hình

    Claude Haiku Latest

    Anthropic🇺🇸
    Phân tích tài liệu#37Toán học#114Phần mềm & Dịch vụ CNTT#116Viết, Văn học & Ngôn ngữ#125

    This is an alias that always points to Anthropic's latest Claude Haiku model.

    Apr 27, 2026·ngữ cảnh 200k·mức chi phí 1x
  • Mô hình

    GPT Mini Latest

    OpenAI🇺🇸
    Hiểu hình ảnh#50Kinh doanh, Quản lý & Tài chính#61Toán học#81Pháp lý & Chính phủ#82

    This is an alias that always points to OpenAI's latest smaller/cheaper GPT model.

    Apr 27, 2026·ngữ cảnh 400k·mức chi phí 0.9x
  • Mô hình

    Gemini Pro Latest

    Google🇺🇸
    Giải trí, Thể thao & Truyền thông#12Viết, Văn học & Ngôn ngữ#13Pháp lý & Chính phủ#14Y học & Chăm sóc sức khỏe#15

    This is an alias that always points to Google's latest Gemini Pro model.

    Apr 27, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Kimi Latest

    Moonshot🇨🇳
    Pháp lý & Chính phủ#3Khoa học Sự sống, Vật lý & Xã hội#5Y học & Chăm sóc sức khỏe#8Toán học#12

    This is an alias that always points to Moonshot AI's latest Kimi model.

    Apr 27, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Gemini Flash Latest

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#9Khoa học Sự sống, Vật lý & Xã hội#10Giải trí, Thể thao & Truyền thông#10Y học & Chăm sóc sức khỏe#10

    This is an alias that always points to Google's latest Gemini Flash model.

    Apr 27, 2026·ngữ cảnh 1M·mức chi phí 0.8x
  • Mô hình

    Claude Sonnet Latest

    Anthropic🇺🇸
    Phần mềm & Dịch vụ CNTT#20Viết, Văn học & Ngôn ngữ#34Pháp lý & Chính phủ#35Hiểu hình ảnh#35

    This is an alias that always points to Anthropic's latest Claude Sonnet model.

    Apr 27, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Qwen3.5 Plus 2026-04-20

    Qwen🇨🇳

    An April 2026 snapshot of Qwen3.5 Plus, Alibaba's proprietary mid-tier model in the Qwen3.5 generation, offering the same general-purpose reasoning and agentic capability as other Plus-tier Qwen3.5 releases.

    Apr 27, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    Qwen3.6 Flash

    Qwen🇨🇳

    Qwen3.6 Flash is Alibaba's fast, low-cost proprietary tier of the Qwen3.6 generation, aimed at high-throughput, latency-sensitive chat and agentic tasks.

    Apr 27, 2026·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    Qwen3.6 35B A3B

    Qwen🇨🇳

    Qwen3.6 35B-A3B is a sparse mixture-of-experts model from Alibaba's Qwen3.6 line (35B total, 3B active parameters), offering efficient general-purpose and agentic capability at low inference cost.

    Apr 27, 2026·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    Qwen3.6 Max Preview

    Qwen🇨🇳
    Toán học#37Y học & Chăm sóc sức khỏe#51Khoa học Sự sống, Vật lý & Xã hội#54Phần mềm & Dịch vụ CNTT#60

    Qwen3.6 Max Preview is a preview of Alibaba's flagship proprietary Qwen3.6 model, intended for top-tier reasoning, coding, and agentic performance ahead of general availability.

    Apr 27, 2026·ngữ cảnh 262k·mức chi phí 1x
  • Mô hình

    Qwen3.6 27B

    Qwen🇨🇳

    Qwen3.6 27B is an open-weight model from Alibaba's Qwen3.6 generation, a successor to Qwen3.5 with incremental improvements to general chat, reasoning, and agentic performance.

    Apr 27, 2026·ngữ cảnh 262k·mức chi phí 0.6x
  • Mô hình

    GPT-5.5 Pro

    OpenAI🇺🇸

    GPT-5.5 Pro is the extended-reasoning tier of GPT-5.5, for the hardest coding, research, and analytical tasks where maximum accuracy is worth the extra latency and cost.

    Apr 24, 2026·ngữ cảnh 1.1M·mức chi phí 35x
  • Mô hình

    GPT-5.5 Pro (batch)

    OpenAI🇺🇸

    GPT-5.5 Pro is the extended-reasoning tier of GPT-5.5 for the hardest coding, research, and analytical tasks. This is the batch variant for cheaper asynchronous processing.

    Apr 24, 2026·ngữ cảnh 1.1M·mức chi phí 18x
  • Mô hình

    GPT-5.5

    OpenAI🇺🇸
    Phân tích tài liệu#8Kinh doanh, Quản lý & Tài chính#10Toán học#14Hiểu hình ảnh#16

    GPT-5.5 is OpenAI's April 2026 flagship model, described as its most intuitive yet at understanding intent and carrying more of a task itself. It excels at writing and debugging code, research, data analysis, and multi-step agentic tool use, with reduced hallucination in sensitive domains like law, medicine, and finance.

    Apr 24, 2026·ngữ cảnh 1.1M·mức chi phí 6x
  • Mô hình

    GPT-5.5 (batch)

    OpenAI🇺🇸
    Phân tích tài liệu#8Kinh doanh, Quản lý & Tài chính#10Toán học#14Hiểu hình ảnh#16

    GPT-5.5 is OpenAI's April 2026 flagship model, excelling at coding, research, data analysis, and agentic tool use with reduced hallucination in sensitive domains. This is the batch variant for cheaper asynchronous processing.

    Apr 24, 2026·ngữ cảnh 1.1M·mức chi phí 3x
  • Mô hình

    DeepSeek V4 Pro 0423

    DeepSeek🇨🇳
    Pháp lý & Chính phủ#37Kinh doanh, Quản lý & Tài chính#44Viết, Văn học & Ngôn ngữ#47Khoa học Sự sống, Vật lý & Xã hội#48

    DeepSeek V4 Pro is the flagship variant of DeepSeek's V4 model family, featuring reasoning-effort levels (low/high/max) and strong performance across coding, reasoning, and agentic benchmarks. It is DeepSeek's top-tier general-purpose and reasoning model in the V4 generation.

    Apr 24, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    DeepSeek V4 Flash 0423

    DeepSeek🇨🇳
    Y học & Chăm sóc sức khỏe#93Pháp lý & Chính phủ#94Khoa học Sự sống, Vật lý & Xã hội#95Viết, Văn học & Ngôn ngữ#99

    DeepSeek V4 Flash is the faster, cheaper variant of DeepSeek's V4 model family, released in public beta and reported to substantially exceed V4-Pro-Preview on agent benchmarks. It is designed for high-throughput agentic and coding workloads where speed and cost matter alongside strong capability.

    Apr 24, 2026·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    Hy3 preview

    Tencent🇨🇳

    Hy3-preview is the early-access preview release of Tencent Hunyuan's Hy3 model, a large MoE architecture with hybrid fast/slow reasoning aimed at coding, research, and agentic workflows. It offers the same general capabilities as the final Hy3 release under a more restrictive preview license.

    Apr 22, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    MiMo-V2.5-Pro

    Xiaomi🇨🇳
    Toán học#30Phần mềm & Dịch vụ CNTT#37Khoa học Sự sống, Vật lý & Xã hội#37Kinh doanh, Quản lý & Tài chính#38

    MiMo v2.5 Pro is the larger, higher-capability variant of Xiaomi's MiMo v2.5 open-source reasoning model line. It targets more demanding reasoning, math, and coding tasks than the base MiMo v2.5 model while retaining an open-source, efficiency-conscious design.

    Apr 22, 2026·ngữ cảnh 1.1M·mức chi phí 0.2x
  • Mô hình

    MiMo-V2.5

    Xiaomi🇨🇳
    Hiểu hình ảnh#65Toán học#76Phần mềm & Dịch vụ CNTT#94Kinh doanh, Quản lý & Tài chính#98

    MiMo v2.5 is part of Xiaomi's MiMo family of open-source reasoning models, following the original MiMo-7B focused on math, code, and reasoning-heavy tasks. It scales up Xiaomi's reasoning-focused training approach for stronger multimodal and general reasoning performance.

    Apr 22, 2026·ngữ cảnh 1.1M·mức chi phí 0.1x
  • Mô hình

    Claude Opus Latest

    Anthropic🇺🇸
    Viết, Văn học & Ngôn ngữ#2Giải trí, Thể thao & Truyền thông#2Y học & Chăm sóc sức khỏe#2Toán học#6

    This is an alias that always points to Anthropic's latest Claude Opus model.

    Apr 21, 2026·ngữ cảnh 1M·mức chi phí 4x
  • Mô hình

    Kimi K2.6

    Moonshot🇨🇳
    Phân tích tài liệu#27Hiểu hình ảnh#38Toán học#42Phần mềm & Dịch vụ CNTT#48

    Kimi K2.6 is Moonshot AI's April 2026 open-weight model that ties GPT-5.5 on coding benchmarks like SWE-Bench Pro while costing roughly 80% less per token. It is designed for high-end coding, reasoning, and agentic tasks at a fraction of proprietary-model cost.

    Apr 20, 2026·ngữ cảnh 262k·mức chi phí 0.8x
  • Mô hình

    Claude Opus 4.7

    Anthropic🇺🇸
    Phần mềm & Dịch vụ CNTT#2Kinh doanh, Quản lý & Tài chính#3Khoa học Sự sống, Vật lý & Xã hội#4Hiểu hình ảnh#4

    Claude Opus 4.7 is the next generation of Anthropic's Opus family, purpose-built for long-running, asynchronous agents that operate with minimal supervision. It carries a 1M-token context window and targets complex, multi-step engineering and research tasks.

    Apr 16, 2026·ngữ cảnh 1M·mức chi phí 5x
  • Mô hình

    Claude Opus 4.7 (batch)

    Anthropic🇺🇸
    Phần mềm & Dịch vụ CNTT#2Kinh doanh, Quản lý & Tài chính#3Khoa học Sự sống, Vật lý & Xã hội#4Hiểu hình ảnh#4

    Batch-processing variant of Claude Opus 4.7, Anthropic's Opus model built for long-running asynchronous agents. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Apr 16, 2026·ngữ cảnh 1M·mức chi phí 2.5x
  • Mô hình

    GLM 5.1

    Z.AI🇨🇳
    Y học & Chăm sóc sức khỏe#45Khoa học Sự sống, Vật lý & Xã hội#47Giải trí, Thể thao & Truyền thông#48Viết, Văn học & Ngôn ngữ#50

    GLM-5.1 is an incremental upgrade to Z.ai's GLM-5 flagship model, released as open source in April 2026 with improved reasoning and agentic performance. It continues GLM-5's focus on complex systems design and large-scale programming tasks.

    Apr 7, 2026·ngữ cảnh 205k·mức chi phí 1x
  • Mô hình

    Gemma 4 26B A4B

    Google🇺🇸
    Toán học#47Hiểu hình ảnh#59Pháp lý & Chính phủ#67Kinh doanh, Quản lý & Tài chính#90

    Gemma 4 26B-A4B IT is a Mixture-of-Experts model from Google's open-weight Gemma 4 family with 26B total but only ~4B active parameters, giving high-throughput inference for advanced reasoning tasks. It is a multimodal, hybrid-thinking model supporting 140+ languages and up to 256K context.

    Apr 3, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Gemma 4 26B A4B (free)

    Google🇺🇸
    Toán học#47Hiểu hình ảnh#59Pháp lý & Chính phủ#67Kinh doanh, Quản lý & Tài chính#90

    Gemma 4 26B-A4B IT is a Mixture-of-Experts model from Google's open-weight Gemma 4 family with 26B total but only ~4B active parameters, giving high-throughput inference for advanced reasoning tasks. This is the free-tier variant of the same model.

    Apr 3, 2026·ngữ cảnh 262k·Miễn phí
  • Mô hình

    Gemma 4 31B

    Google🇺🇸
    Phân tích tài liệu#35Hiểu hình ảnh#42Toán học#53Phần mềm & Dịch vụ CNTT#68

    Gemma 4 31B IT is a dense model from Google's open-weight Gemma 4 family that bridges server-grade performance and local execution, ranking near the top of open-model leaderboards. It handles multimodal, hybrid-thinking reasoning across 140+ languages with up to 256K context, well suited for advanced chat, coding, and reasoning.

    Apr 2, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Gemma 4 31B (free)

    Google🇺🇸
    Phân tích tài liệu#35Hiểu hình ảnh#42Toán học#53Phần mềm & Dịch vụ CNTT#68

    Gemma 4 31B IT is a dense model from Google's open-weight Gemma 4 family that bridges server-grade performance and local execution, ranking near the top of open-model leaderboards. This is the free-tier variant of the same model.

    Apr 2, 2026·ngữ cảnh 262k·Miễn phí
  • Mô hình

    Qwen3.6 Plus

    Qwen🇨🇳
    Toán học#73Kinh doanh, Quản lý & Tài chính#87Phần mềm & Dịch vụ CNTT#88Viết, Văn học & Ngôn ngữ#89

    Qwen3.6 Plus is Alibaba's proprietary mid-tier model in the Qwen3.6 generation, balancing cost and capability for general chat, reasoning, and agentic use cases.

    Apr 2, 2026·ngữ cảnh 1M·mức chi phí 0.4x
  • Mô hình

    GLM 5V Turbo

    Z.AI🇨🇳
    Phân tích tài liệu#39Hiểu hình ảnh#67Toán học#94Phần mềm & Dịch vụ CNTT#101

    GLM-5V-Turbo is a fast, lower-latency vision-language variant in Z.ai's GLM-5 family, combining multimodal image/video understanding with efficient serving. It is designed for agentic and multimodal workflows where speed and cost matter alongside visual reasoning.

    Apr 1, 2026·ngữ cảnh 203k·mức chi phí 0.9x
  • Mô hình

    Trinity Large Thinking

    Arcee AI🇺🇸
    Phần mềm & Dịch vụ CNTT#171Giải trí, Thể thao & Truyền thông#171Y học & Chăm sóc sức khỏe#173Khoa học Sự sống, Vật lý & Xã hội#175

    Trinity-Large-Thinking is Arcee AI's open-source frontier reasoning model, a 399B-parameter sparse mixture-of-experts model (about 13B active per token) with a 512K-token context window, released under Apache 2.0. It generates explicit chain-of-thought reasoning traces and is built for complex, long-horizon agentic tasks and long-document analysis.

    Apr 1, 2026·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    Grok 4.20 Multi-Agent

    xAI🇺🇸

    Grok 4.20 Multi-Agent is a variant of xAI's Grok 4.20 that runs several specialized agents on a shared backbone within a single API call to perform deep, multi-step research. It combines web search, code execution, and data synthesis tools to deliver comprehensive, well-sourced answers for research-heavy tasks.

    Mar 31, 2026·ngữ cảnh 2M·mức chi phí 0.6x
  • Mô hình

    Grok 4.20

    xAI🇺🇸

    Grok 4.20 is an xAI flagship model built for long-running agentic work and ambitious interactive and visual tasks, with a 500,000-token context window and text/image input. It targets high-intelligence coding, reasoning, and multi-step agent use cases.

    Mar 31, 2026·ngữ cảnh 2M·mức chi phí 0.6x
  • Mô hình

    Lyria 3 Pro Preview

    Google🇺🇸

    Lyria 3 Pro is Google DeepMind's flagship music generation model, capable of producing full compositions up to three minutes long with structural awareness of intros, verses, choruses, and bridges. It supports multi-language vocals and genre versatility for professional-grade, prompt-driven music creation.

    Mar 30, 2026·ngữ cảnh 1M·Miễn phí
  • Mô hình

    Lyria 3 Clip Preview

    Google🇺🇸

    Lyria 3 (clip preview) is Google DeepMind's music generation model that produces high-fidelity stereo audio clips up to 30 seconds from text prompts. It is designed for rapid prototyping, social media assets, and short-form audio generation rather than full song composition.

    Mar 30, 2026·ngữ cảnh 1M·Miễn phí
  • Mô hình

    Reka Edge

    Reka AI🇺🇸

    Reka Edge is a small, efficient model from Reka AI designed for low-latency, on-device or resource-constrained deployments. It targets simple chat and task completion where speed and low cost outweigh the need for top-end capability.

    Mar 20, 2026·ngữ cảnh 16k·mức chi phí 0.1x
  • Mô hình

    MiniMax M2.7

    Minimax🇨🇳
    Toán học#101Phần mềm & Dịch vụ CNTT#111Kinh doanh, Quản lý & Tài chính#117Khoa học Sự sống, Vật lý & Xã hội#137

    MiniMax-M2.7 is MiniMax's model for software engineering, agent workflows, and professional document tasks, notable as a self-evolving model that ran autonomous reinforcement-learning optimization loops during training. It targets advanced coding and long-horizon agentic use cases.

    Mar 18, 2026·ngữ cảnh 205k·mức chi phí 0.2x
  • Mô hình

    GPT-5.4 Nano

    OpenAI🇺🇸
    Hiểu hình ảnh#87Toán học#119Phần mềm & Dịch vụ CNTT#139Kinh doanh, Quản lý & Tài chính#151

    GPT-5.4 Nano is the smallest, fastest model in the GPT-5.4 family, aimed at high-throughput, low-latency tasks like classification and simple extraction where cost efficiency is the priority.

    Mar 17, 2026·ngữ cảnh 400k·mức chi phí 0.2x
  • Mô hình

    GPT-5.4 Nano (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#87Toán học#119Phần mềm & Dịch vụ CNTT#139Kinh doanh, Quản lý & Tài chính#151

    GPT-5.4 Nano is the smallest, fastest model in the GPT-5.4 family, aimed at high-throughput, low-latency tasks. This is the batch variant for cheaper asynchronous processing.

    Mar 17, 2026·ngữ cảnh 400k·mức chi phí 0.1x
  • Mô hình

    GPT-5.4 Mini

    OpenAI🇺🇸
    Hiểu hình ảnh#50Kinh doanh, Quản lý & Tài chính#61Toán học#81Pháp lý & Chính phủ#82

    GPT-5.4 Mini is the smaller, cheaper counterpart to GPT-5.4, suited to everyday chat, lightweight coding, and agentic tasks where speed and cost outweigh maximum reasoning depth.

    Mar 17, 2026·ngữ cảnh 400k·mức chi phí 0.9x
  • Mô hình

    GPT-5.4 Mini (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#50Kinh doanh, Quản lý & Tài chính#61Toán học#81Pháp lý & Chính phủ#82

    GPT-5.4 Mini is the smaller, cheaper counterpart to GPT-5.4 for everyday chat, lightweight coding, and agentic tasks. This is the batch variant for cheaper asynchronous processing.

    Mar 17, 2026·ngữ cảnh 400k·mức chi phí 0.4x
  • Mô hình

    Mistral Small 4

    Mistral🇫🇷
    Hiểu hình ảnh#117Y học & Chăm sóc sức khỏe#194Phần mềm & Dịch vụ CNTT#202Kinh doanh, Quản lý & Tài chính#203

    Mistral Small 4 (2603) is Mistral AI's March 2026 model, a 119B-parameter mixture-of-experts model with 6B active parameters that unifies instruct, reasoning, and coding with a toggleable reasoning mode and 256K context. It combines reasoning strength from Magistral, vision from Pixtral, and coding ability from Devstral in one efficient model.

    Mar 16, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Mistral Small 4 (batch)

    Mistral🇫🇷
    Hiểu hình ảnh#117Y học & Chăm sóc sức khỏe#194Phần mềm & Dịch vụ CNTT#202Kinh doanh, Quản lý & Tài chính#203

    This is the batch-processing variant of Mistral Small 4 (mistral-small-2603), Mistral AI's 119B-parameter mixture-of-experts model (6.5B active, Apache 2.0 license, 262K-token context) that unifies reasoning, vision, and agentic coding into one model. Positioned as an efficient mid-tier workhorse for complex analysis, coding, and multimodal tasks, the batch endpoint processes requests asynchronously at a lower price for high-volume, non-latency-sensitive workloads.

    Mar 16, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GLM 5 Turbo

    Z.AI🇨🇳

    GLM-5-Turbo is a faster, lower-latency serving variant of Z.ai's flagship GLM-5 model, aimed at reducing cost and response time while retaining GLM-5's strengths in coding and long-horizon agentic workflows.

    Mar 15, 2026·ngữ cảnh 203k·mức chi phí 0.9x
  • Mô hình

    Nemotron 3 Super

    NVIDIA🇺🇸

    Nemotron 3 Super is a 120B mixture-of-experts model with 12B active parameters in NVIDIA's Nemotron 3 family, sitting between Nano and Ultra in scale. It offers stronger reasoning and agentic capability than Nemotron 3 Nano while remaining more efficient than Nemotron 3 Ultra.

    Mar 11, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Nemotron 3 Super (free)

    NVIDIA🇺🇸

    Nemotron 3 Super is a 120B mixture-of-experts model with 12B active parameters in NVIDIA's Nemotron 3 family, offering stronger reasoning and agentic capability than Nemotron 3 Nano while staying more efficient than Nemotron 3 Ultra. This is the free-tier variant of the same model.

    Mar 11, 2026·ngữ cảnh 262k·Miễn phí
  • Mô hình

    Seed-2.0-Lite

    Bytedance🇨🇳

    Seed 2.0 Lite is a cost-efficient, low-latency multimodal model from ByteDance Seed, delivering solid text, vision, and agentic tool-use capability. It is designed as a practical default for high-frequency production workloads.

    Mar 10, 2026·ngữ cảnh 262k·mức chi phí 0.4x
  • Mô hình

    Qwen3.5-9B

    Qwen🇨🇳

    Qwen3.5 9B is a small open-weight model from Alibaba's Qwen3.5 generation, suited for fast, low-cost general chat and simple coding tasks with the efficiency gains of the 3.5 series.

    Mar 10, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GPT-5.4 Pro

    OpenAI🇺🇸

    GPT-5.4 Pro is the extended-reasoning tier of GPT-5.4, prized for low hallucination rates on hard problems. It suits complex research, coding, and analysis where reliability matters more than speed.

    Mar 5, 2026·ngữ cảnh 1.1M·mức chi phí 35x
  • Mô hình

    GPT-5.4 Pro (batch)

    OpenAI🇺🇸

    GPT-5.4 Pro is the extended-reasoning tier of GPT-5.4 for complex research, coding, and analysis that need reliable, low-hallucination answers. This is the batch variant for cheaper asynchronous processing.

    Mar 5, 2026·ngữ cảnh 1.1M·mức chi phí 18x
  • Mô hình

    GPT-5.4

    OpenAI🇺🇸
    Phân tích tài liệu#14Toán học#19Hiểu hình ảnh#22Kinh doanh, Quản lý & Tài chính#25

    GPT-5.4 is OpenAI's March 2026 flagship model, released in Thinking and Pro variants known for low hallucination rates. It targets advanced reasoning, research, and coding tasks that need reliable, well-grounded answers.

    Mar 5, 2026·ngữ cảnh 1.1M·mức chi phí 3x
  • Mô hình

    GPT-5.4 (batch)

    OpenAI🇺🇸
    Phân tích tài liệu#14Toán học#19Hiểu hình ảnh#22Kinh doanh, Quản lý & Tài chính#25

    GPT-5.4 is OpenAI's March 2026 flagship model known for low hallucination rates, suited to advanced reasoning, research, and coding. This is the batch variant for cheaper asynchronous processing.

    Mar 5, 2026·ngữ cảnh 1.1M·mức chi phí 1.5x
  • Mô hình

    Mercury 2

    Inception🇺🇸
    Phần mềm & Dịch vụ CNTT#219Giải trí, Thể thao & Truyền thông#219Kinh doanh, Quản lý & Tài chính#227Y học & Chăm sóc sức khỏe#242

    Mercury 2 is Inception Labs' diffusion-based LLM that generates responses in parallel rather than token-by-token, delivering very high throughput (1000+ tokens/second) with a 128K context window. It offers tunable reasoning and native tool use, matching Claude Haiku/Gemini Flash-class quality at much faster speeds and lower cost, ideal for latency-sensitive agentic and coding tasks.

    Mar 4, 2026·ngữ cảnh 128k·mức chi phí 0.2x
  • Mô hình

    Gemini 3.1 Flash Lite Preview

    Google🇺🇸
    Hiểu hình ảnh#64Viết, Văn học & Ngôn ngữ#94Giải trí, Thể thao & Truyền thông#102Y học & Chăm sóc sức khỏe#102

    Gemini 3.1 Flash-Lite Preview is the preview release of Google's high-efficiency multimodal model, optimized for low-latency, high-volume workloads such as lightweight agentic workflows and simple data extraction. It offers the same responsiveness and cost focus as the GA Flash-Lite release.

    Mar 3, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    Seed-2.0-Mini

    Bytedance🇨🇳

    Seed 2.0 Mini is ByteDance Seed's smallest 2.0-generation model, targeting latency-sensitive, high-concurrency, cost-sensitive scenarios with a 256K context window and four reasoning-effort modes. It delivers performance comparable to Seed 1.6 at lower cost.

    Feb 26, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Qwen3.5-35B-A3B

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#156Kinh doanh, Quản lý & Tài chính#159Toán học#160Khoa học Sự sống, Vật lý & Xã hội#161

    Qwen3.5 35B-A3B is a sparse mixture-of-experts model from Alibaba's Qwen3.5 line (35B total, 3B active parameters), built for efficient general-purpose and agentic tasks at low inference cost.

    Feb 25, 2026·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    Qwen3.5-27B

    Qwen🇨🇳
    Hiểu hình ảnh#73Khoa học Sự sống, Vật lý & Xã hội#124Toán học#124Phần mềm & Dịch vụ CNTT#143

    Qwen3.5 27B is a mid-size dense open-weight model from Alibaba's Qwen3.5 generation, offering general-purpose chat, coding, and agentic capability with efficiency improvements over the Qwen3 series.

    Feb 25, 2026·ngữ cảnh 262k·mức chi phí 0.3x
  • Mô hình

    Qwen3.5-122B-A10B

    Qwen🇨🇳
    Hiểu hình ảnh#69Kinh doanh, Quản lý & Tài chính#122Toán học#125Khoa học Sự sống, Vật lý & Xã hội#133

    Qwen3.5 122B-A10B is one of Alibaba's Qwen3.5 open-weight mixture-of-experts models (122B total, 10B active), part of a generation positioned for the 'agentic AI era' with improved tool use and multi-step task handling over Qwen3.

    Feb 25, 2026·ngữ cảnh 262k·mức chi phí 0.4x
  • Mô hình

    Qwen3.5-Flash

    Qwen🇨🇳
    Khoa học Sự sống, Vật lý & Xã hội#154Toán học#156Kinh doanh, Quản lý & Tài chính#157Phần mềm & Dịch vụ CNTT#163

    Qwen3.5 Flash is Alibaba's fast, low-cost proprietary tier of the Qwen3.5 generation (this snapshot dated February 2026), designed for high-throughput, latency-sensitive general and agentic tasks.

    Feb 25, 2026·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Gemini 3.1 Pro Preview Custom Tools

    Google🇺🇸

    Gemini 3.1 Pro Preview (custom tools) is a variant of Google's Gemini 3.1 Pro Preview configured for custom tool-calling setups, sharing the base model's strengths in agentic coding, long-horizon tool orchestration, structured planning, and multimodal analysis.

    Feb 25, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    GPT-5.3-Codex

    OpenAI🇺🇸

    GPT-5.3-Codex is OpenAI's February 2026 Codex-native coding agent, pairing frontier coding performance with general reasoning for long-horizon, real-world engineering work. It supports steering the agent mid-task and is notable as a model OpenAI used to help build itself.

    Feb 24, 2026·ngữ cảnh 400k·mức chi phí 2.5x
  • Mô hình

    Aion-2.0

    Aion Labs🇵🇱

    Aion-2.0 is a variant of DeepSeek V3.2 from AionLabs, tuned specifically for immersive roleplaying and collaborative storytelling. It is particularly strong at introducing tension, conflict, and mature or darker themes with nuance and depth.

    Feb 23, 2026·ngữ cảnh 131k·mức chi phí 0.4x
  • Mô hình

    Gemini 3.1 Pro Preview

    Google🇺🇸
    Giải trí, Thể thao & Truyền thông#12Viết, Văn học & Ngôn ngữ#13Pháp lý & Chính phủ#14Y học & Chăm sóc sức khỏe#15

    Gemini 3.1 Pro Preview is Google's advanced reasoning model improving long-horizon stability and tool orchestration over Gemini 3 Pro, with a new medium thinking level to balance cost, speed, and performance. It excels at agentic coding, structured planning, multimodal analysis, and workflow automation for autonomous agents and enterprise tasks.

    Feb 19, 2026·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Gemini 3.1 Pro Preview (batch)

    Google🇺🇸
    Giải trí, Thể thao & Truyền thông#12Viết, Văn học & Ngôn ngữ#13Pháp lý & Chính phủ#14Y học & Chăm sóc sức khỏe#15

    Gemini 3.1 Pro Preview is Google's advanced reasoning model improving long-horizon stability and tool orchestration, well suited to agentic coding, structured planning, multimodal analysis, and workflow automation. This is the batch variant for asynchronous, lower-cost processing.

    Feb 19, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    Claude Sonnet 4.6

    Anthropic🇺🇸
    Phân tích tài liệu#11Khoa học Sự sống, Vật lý & Xã hội#22Kinh doanh, Quản lý & Tài chính#24Phần mềm & Dịch vụ CNTT#25

    Claude Sonnet 4.6 is a further refined mid-tier model, engineered for complex enterprise workflows, deep software engineering, and autonomous agentic execution with a 1M-token context window in beta. It scores strongly on coding (SWE-bench) and computer-use benchmarks while keeping Sonnet-tier pricing.

    Feb 17, 2026·ngữ cảnh 1M·mức chi phí 3x
  • Mô hình

    Claude Sonnet 4.6 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#11Khoa học Sự sống, Vật lý & Xã hội#22Kinh doanh, Quản lý & Tài chính#24Phần mềm & Dịch vụ CNTT#25

    Batch-processing variant of Claude Sonnet 4.6, Anthropic's mid-tier model for enterprise coding and agentic execution. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Feb 17, 2026·ngữ cảnh 1M·mức chi phí 1.5x
  • Mô hình

    Qwen3.5 Plus 2026-02-15

    Qwen🇨🇳

    A February 2026 snapshot of Qwen3.5 Plus, Alibaba's proprietary mid-tier Qwen3.5 model, balancing cost and capability for general chat, reasoning, and agentic workloads.

    Feb 16, 2026·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    Qwen3.5 397B A17B

    Qwen🇨🇳
    Hiểu hình ảnh#54Toán học#75Y học & Chăm sóc sức khỏe#81Khoa học Sự sống, Vật lý & Xã hội#85

    Qwen3.5 397B-A17B is the largest and first-released model in Alibaba's Qwen3.5 mixture-of-experts generation, targeting flagship-level reasoning, coding, and agentic performance among open-weight models.

    Feb 16, 2026·ngữ cảnh 262k·mức chi phí 0.7x
  • Mô hình

    MiniMax M2.5

    Minimax🇨🇳
    Phần mềm & Dịch vụ CNTT#157Pháp lý & Chính phủ#161Kinh doanh, Quản lý & Tài chính#162Giải trí, Thể thao & Truyền thông#164

    MiniMax-M2.5 is a further update in MiniMax's M2 series, achieving state-of-the-art results among MiniMax models in programming, tool calling, search, and office productivity tasks. It is aimed at real-world developer and knowledge-work productivity.

    Feb 12, 2026·ngữ cảnh 205k·mức chi phí 0.2x
  • Mô hình

    GLM 5

    Z.AI🇨🇳
    Giải trí, Thể thao & Truyền thông#53Viết, Văn học & Ngôn ngữ#55Khoa học Sự sống, Vật lý & Xã hội#60Pháp lý & Chính phủ#60

    GLM-5 is Z.ai's flagship open-source foundation model, scaling to 744B total parameters (40B active) and trained on 28.5T tokens. It is engineered for complex systems design and long-horizon agent workflows, delivering production-grade performance on large-scale programming tasks for expert developers.

    Feb 11, 2026·ngữ cảnh 205k·mức chi phí 0.4x
  • Mô hình

    Qwen3 Max Thinking

    Qwen🇨🇳

    The extended chain-of-thought variant of Qwen3 Max, Alibaba's flagship model, built for the hardest reasoning, math, and multi-step agentic tasks where deliberate step-by-step thinking improves accuracy.

    Feb 9, 2026·ngữ cảnh 262k·mức chi phí 0.8x
  • Mô hình

    Claude Opus 4.6

    Anthropic🇺🇸
    Khoa học Sự sống, Vật lý & Xã hội#2Phần mềm & Dịch vụ CNTT#3Toán học#3Hiểu hình ảnh#3

    Claude Opus 4.6 is Anthropic's strongest model for coding and long-running professional tasks in this generation, with a 1M-token context window and 128K max output. It is designed for the most demanding software engineering and multi-step agentic workflows.

    Feb 4, 2026·ngữ cảnh 1M·mức chi phí 5x
  • Mô hình

    Claude Opus 4.6 (batch)

    Anthropic🇺🇸
    Khoa học Sự sống, Vật lý & Xã hội#2Phần mềm & Dịch vụ CNTT#3Toán học#3Hiểu hình ảnh#3

    Batch-processing variant of Claude Opus 4.6, Anthropic's strongest coding-focused Opus model for long-running professional tasks. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Feb 4, 2026·ngữ cảnh 1M·mức chi phí 2.5x
  • Mô hình

    Qwen3 Coder Next

    Qwen🇨🇳

    Qwen3 Coder Next is a next-generation iteration of Alibaba's Qwen3 Coder line, likely built on the more efficient Qwen3-Next architecture, aimed at agentic coding with improved efficiency and long-context repository handling.

    Feb 4, 2026·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    Free Models Router

    OpenRouter🇺🇸

    OpenRouter Free is a meta-routing model from OpenRouter that automatically selects among available free-tier models to answer a request. It suits experimentation, low-stakes tasks, and cost-free testing rather than production workloads needing a specific model's guarantees.

    Feb 1, 2026·ngữ cảnh 200k·Miễn phí
  • Mô hình

    Step 3.5 Flash

    Stepfun🇨🇳
    Toán học#132Phần mềm & Dịch vụ CNTT#156Kinh doanh, Quản lý & Tài chính#170Viết, Văn học & Ngôn ngữ#176

    Step 3.5 Flash is StepFun's flagship sparse MoE reasoning model (196B parameters, 11B active) built for fast, reliable execution of complex tasks. It excels at decomposing and planning multi-step problems and orchestrating tool calls, making it well suited for agentic and coding workflows.

    Jan 29, 2026·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Kimi K2.5

    Moonshot🇨🇳
    Phân tích tài liệu#34Toán học#48Hiểu hình ảnh#48Phần mềm & Dịch vụ CNTT#77

    Kimi K2.5 is Moonshot AI's January 2026 update to Kimi K2, a 1T-parameter mixture-of-experts model with 32B active parameters and multimodal vision-language understanding, offering both instant and thinking modes. It is built for advanced agentic and conversational workflows.

    Jan 27, 2026·ngữ cảnh 262k·mức chi phí 0.5x
  • Mô hình

    Solar Pro 3

    Upstage🇰🇷

    Solar Pro 3 is Upstage's efficient large language model, building on Solar Pro 2 with Upstage's SnapPO reinforcement learning framework to sharpen step-by-step reasoning in math, code, and agentic tasks. It roughly doubles agentic benchmark performance over its predecessor while keeping the same API and serving footprint, with strong multilingual (notably Korean) support.

    Jan 27, 2026·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    MiniMax M2-her

    Minimax🇨🇳

    MiniMax-M2-her is a dialogue-first model built for immersive roleplay, character-driven chat, and expressive multi-turn conversation, distilled from years of MiniMax's role-play optimization work. It powers MiniMax's Talkie companion app and is suited to storytelling and companion-style interactions rather than coding or agentic tasks.

    Jan 23, 2026·ngữ cảnh 66k·mức chi phí 0.3x
  • Palmyra X5

    writer

    Palmyra X5 is Writer's most advanced enterprise foundation model, built for building and scaling AI agents with a 1M-token context window and a hybrid attention architecture. It offers advanced reasoning, multi-step tool-calling, built-in RAG, and multilingual support tailored for enterprise workflows.

    Jan 21, 2026·ngữ cảnh 1M·mức chi phí 1x
  • Mô hình

    GPT Audio

    OpenAI🇺🇸

    GPT-Audio is OpenAI's model for natural speech input and output, generating expressive spoken responses directly from audio prompts. It is designed for voice assistants, spoken dialogue apps, and other audio-native interactions.

    Jan 19, 2026·ngữ cảnh 128k·mức chi phí 2x
  • Mô hình

    GPT Audio Mini

    OpenAI🇺🇸

    GPT-Audio Mini is a smaller, cheaper version of OpenAI's GPT-Audio model for speech input and output, suited to cost-sensitive or high-volume voice assistant and spoken dialogue applications.

    Jan 19, 2026·ngữ cảnh 128k·mức chi phí 0.5x
  • Mô hình

    GLM 4.7 Flash

    Z.AI🇨🇳
    Phần mềm & Dịch vụ CNTT#181Toán học#185Kinh doanh, Quản lý & Tài chính#194Pháp lý & Chính phủ#199

    GLM-4.7-Flash is a 30B-class SOTA model from Z.ai that balances performance and efficiency, offering a faster, cheaper alternative to the full GLM-4.7 model. It targets general reasoning and coding tasks where speed and cost matter more than maximum capability.

    Jan 19, 2026·ngữ cảnh 200k·mức chi phí 0.1x
  • Mô hình

    GPT-5.2-Codex

    OpenAI🇺🇸

    GPT-5.2-Codex is the coding-specialized variant of GPT-5.2, built for agentic software engineering across large codebases. It benefits from GPT-5.2's large context window for repository-scale refactors and long-horizon coding tasks.

    Jan 14, 2026·ngữ cảnh 400k·mức chi phí 2.5x
  • Mô hình

    Seed 1.6 Flash

    Bytedance🇨🇳

    Seed 1.6 Flash is a faster, lower-latency variant of ByteDance Seed's Seed 1.6 multimodal model, trading some depth for speed and cost efficiency. It suits high-throughput text and vision tasks where responsiveness matters most.

    Dec 23, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Seed 1.6

    Bytedance🇨🇳

    Seed 1.6 is ByteDance Seed's general-purpose multimodal model with adaptive deep-thinking capability and a 256K-token context window. It is designed as a versatile foundation model for text, vision, and reasoning tasks.

    Dec 23, 2025·ngữ cảnh 262k·mức chi phí 0.4x
  • Mô hình

    MiniMax M2.1

    Minimax🇨🇳

    MiniMax-M2.1 is an update to MiniMax's M2 coding and agentic model line, focused on polyglot programming mastery and precision code refactoring. It targets developers needing strong multi-language coding assistance and reliable agentic tool use.

    Dec 23, 2025·ngữ cảnh 205k·mức chi phí 0.3x
  • Mô hình

    GLM 4.7

    Z.AI🇨🇳
    Y học & Chăm sóc sức khỏe#37Khoa học Sự sống, Vật lý & Xã hội#70Pháp lý & Chính phủ#89Kinh doanh, Quản lý & Tài chính#95

    GLM-4.7 is Z.ai's open-source model generation released after GLM-4.6, continuing incremental improvements in reasoning and coding. It serves as the base for the GLM-4.7-Flash variant that balances performance and efficiency at a smaller 30B class size.

    Dec 22, 2025·ngữ cảnh 205k·mức chi phí 0.5x
  • Mô hình

    Gemini 3 Flash Preview

    Google🇺🇸

    Gemini 3 Flash Preview is Google's fast workhorse model in the Gemini 3 line, tuned for agentic workflows, coding, and complex multi-step reasoning at low latency. It became the default model in the Gemini app for everyday responsive tasks.

    Dec 17, 2025·ngữ cảnh 1M·mức chi phí 0.6x
  • Mô hình

    Gemini 3 Flash Preview (batch)

    Google🇺🇸

    Gemini 3 Flash Preview is Google's fast workhorse model in the Gemini 3 line, tuned for agentic workflows, coding, and complex multi-step reasoning at low latency. This is the batch variant for asynchronous, lower-cost processing.

    Dec 17, 2025·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    Nemotron 3 Nano 30B A3B

    NVIDIA🇺🇸

    Nemotron 3 Nano is NVIDIA's compute-efficient open model in the Nemotron 3 family, a 30B mixture-of-experts model with 3B active parameters. It is optimized for high-throughput, low-cost agentic and multi-agent workloads.

    Dec 14, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GPT-5.2 Chat

    OpenAI🇺🇸
    Kinh doanh, Quản lý & Tài chính#18Y học & Chăm sóc sức khỏe#21Hiểu hình ảnh#31Phần mềm & Dịch vụ CNTT#34

    GPT-5.2 Chat is the fast, conversational (instant) variant of GPT-5.2, tuned for responsive everyday chat rather than deep extended reasoning. It suits general assistant use cases where low latency matters most.

    Dec 10, 2025·ngữ cảnh 128k·mức chi phí 2.5x
  • Mô hình

    GPT-5.2 Pro

    OpenAI🇺🇸

    GPT-5.2 Pro is the extended-reasoning tier of GPT-5.2, trading latency for maximum accuracy on hard problems. It is aimed at complex research, coding, and analytical work with OpenAI's largest available context window.

    Dec 10, 2025·ngữ cảnh 400k·mức chi phí 32x
  • Mô hình

    GPT-5.2 Pro (batch)

    OpenAI🇺🇸

    GPT-5.2 Pro is the extended-reasoning tier of GPT-5.2 for complex research, coding, and analytical work. This is the batch variant for cheaper asynchronous processing.

    Dec 10, 2025·ngữ cảnh 400k·mức chi phí 16x
  • Mô hình

    GPT-5.2

    OpenAI🇺🇸
    Kinh doanh, Quản lý & Tài chính#18Y học & Chăm sóc sức khỏe#21Hiểu hình ảnh#31Phần mềm & Dịch vụ CNTT#34

    GPT-5.2 is OpenAI's December 2025 flagship release, offered in instant, thinking, and pro modes with a 400K-token context window. It is designed for enterprise-scale reasoning, document and codebase analysis, and complex agentic workflows.

    Dec 10, 2025·ngữ cảnh 400k·mức chi phí 2.5x
  • Mô hình

    GPT-5.2 (batch)

    OpenAI🇺🇸
    Kinh doanh, Quản lý & Tài chính#18Y học & Chăm sóc sức khỏe#21Hiểu hình ảnh#31Phần mềm & Dịch vụ CNTT#34

    GPT-5.2 is OpenAI's December 2025 flagship release with a 400K-token context window, suited to enterprise-scale reasoning, document analysis, and agentic workflows. This is the batch variant for cheaper asynchronous processing.

    Dec 10, 2025·ngữ cảnh 400k·mức chi phí 1x
  • Mô hình

    Devstral 2 2512

    Mistral🇫🇷

    Devstral 2 2512 is Mistral AI's open-weight agentic coding model (123B-parameter dense transformer, 256K-token context), released December 2025 and built for tool use that explores codebases, edits multiple files, and powers software-engineering agents. It scores 72.2% on SWE-Bench Verified and 61.3% on SWE-Bench Multilingual, positioning it as a top open-source model for project-level coding and long-horizon agent workflows.

    Dec 9, 2025·ngữ cảnh 262k·mức chi phí 0.4x
  • Mô hình

    Relace Search

    Relace🇺🇸

    Relace Search is a companion model to Relace's code-apply system, built for fast retrieval of relevant code context and locations within a codebase. It is designed for AI coding agents that need quick, accurate code search before generating or applying edits.

    Dec 8, 2025·ngữ cảnh 256k·mức chi phí 0.7x
  • Mô hình

    GLM 4.6V

    Z.AI🇨🇳
    Hiểu hình ảnh#103Toán học#107Y học & Chăm sóc sức khỏe#111Giải trí, Thể thao & Truyền thông#117

    GLM-4.6V is Z.ai's vision-language model that treats images, video, and tools as first-class agent inputs, extending training context to 128K tokens with native multimodal function calling. It is designed for agentic multimodal workflows such as screenshot- and document-driven tool use.

    Dec 8, 2025·ngữ cảnh 131k·mức chi phí 0.2x
  • Mô hình

    GPT-5.1-Codex-Max

    OpenAI🇺🇸

    GPT-5.1-Codex-Max is OpenAI's frontier agentic coding model, using a compaction technique to coherently work across millions of tokens in a single task. It targets project-scale refactors, extended debugging sessions, and multi-hour autonomous coding loops.

    Dec 4, 2025·ngữ cảnh 400k·mức chi phí 2x
  • Mô hình

    Nova 2 Lite

    AWS🇺🇸
    Kinh doanh, Quản lý & Tài chính#211Phần mềm & Dịch vụ CNTT#214Toán học#219Y học & Chăm sóc sức khỏe#220

    Amazon Nova 2 Lite is a fast, cost-effective multimodal reasoning model from AWS that handles text, image, video, and document input with a 1M-token context window. It supports optional extended thinking with adjustable reasoning budgets plus built-in web grounding and code interpreter tools, aimed at everyday production workloads that need quick, affordable responses.

    Dec 2, 2025·ngữ cảnh 1M·mức chi phí 0.5x
  • Mô hình

    Ministral 3 14B 2512

    Mistral🇫🇷

    Ministral 3 14B is Mistral AI's December 2025 edge-focused model, offered in base, instruct, and reasoning variants with a 256K context window. It targets on-device and edge deployment with function calling and multilingual support, offering the strongest quality of the Ministral 3 sizes.

    Dec 2, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Ministral 3 8B 2512

    Mistral🇫🇷
    Y học & Chăm sóc sức khỏe#312Pháp lý & Chính phủ#314Phần mềm & Dịch vụ CNTT#316Kinh doanh, Quản lý & Tài chính#317

    Ministral 3 8B is the mid-sized model in Mistral AI's December 2025 Ministral 3 edge family, balancing capability and efficiency for on-device deployment. It supports a 256K context window, function calling, and multilingual chat.

    Dec 2, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Ministral 3 8B 2512 (batch)

    Mistral🇫🇷
    Y học & Chăm sóc sức khỏe#312Pháp lý & Chính phủ#314Phần mềm & Dịch vụ CNTT#316Kinh doanh, Quản lý & Tài chính#317

    This is the batch-processing variant of Ministral 3 8B (ministral-8b-2512), Mistral AI's dense 8.4B-parameter model plus a 0.4B vision encoder (Apache 2.0 license, 262K-token context, multilingual and multimodal), released December 2025. Designed as a small, efficient edge-capable model that fits in 12GB of VRAM, it suits lightweight and on-device deployments, and the batch endpoint adds a further cost reduction for high-volume asynchronous processing.

    Dec 2, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Ministral 3 3B 2512

    Mistral🇫🇷

    Ministral 3 3B is the smallest model in Mistral AI's December 2025 Ministral 3 edge family, designed for lightweight on-device deployment. It trades some capability for a tiny footprint, suiting simple chat, classification, and function-calling tasks on constrained hardware.

    Dec 2, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Mistral Large 3 2512

    Mistral🇫🇷
    Giải trí, Thể thao & Truyền thông#240Pháp lý & Chính phủ#245Viết, Văn học & Ngôn ngữ#252Toán học#263

    Mistral Large 3 (2512) is Mistral AI's December 2025 flagship model, a sparse mixture-of-experts architecture with 41B active of 675B total parameters, released under Apache 2.0. It is Mistral's most capable model to date for advanced reasoning, coding, and multilingual enterprise workloads.

    Dec 1, 2025·ngữ cảnh 262k·mức chi phí 0.3x
  • Mô hình

    Mistral Large 3 2512 (batch)

    Mistral🇫🇷
    Giải trí, Thể thao & Truyền thông#240Pháp lý & Chính phủ#245Viết, Văn học & Ngôn ngữ#252Toán học#263

    This is the batch-processing variant of Mistral Large 3 (mistral-large-2512), Mistral AI's 675B-parameter sparse mixture-of-experts flagship (41B active, Apache 2.0 license, 262K-token context) released December 2025. It targets complex reasoning, agentic, and long-context workloads, and the batch endpoint offers the same capability at a reduced price via asynchronous processing, ideal for large-scale, non-real-time jobs.

    Dec 1, 2025·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    DeepSeek V3.2

    DeepSeek🇨🇳
    Toán học#85Viết, Văn học & Ngôn ngữ#106Khoa học Sự sống, Vật lý & Xã hội#108Phần mềm & Dịch vụ CNTT#115

    DeepSeek V3.2 is DeepSeek's V3-generation model upgrade underlying both the deepseek-chat (non-thinking) and deepseek-reasoner (thinking) modes, improving efficiency and reasoning quality over V3.1. It is designed as a strong, low-cost general-purpose and reasoning model.

    Dec 1, 2025·ngữ cảnh 164k·mức chi phí 0.1x
  • Mô hình

    Claude Opus 4.5

    Anthropic🇺🇸
    Phân tích tài liệu#21Viết, Văn học & Ngôn ngữ#27Giải trí, Thể thao & Truyền thông#27Pháp lý & Chính phủ#34

    Claude Opus 4.5 is a further refinement of Anthropic's Opus flagship line, with improved coding, reasoning, and long-horizon agentic capabilities over Opus 4.1. It targets the most demanding professional software engineering and analysis tasks.

    Nov 24, 2025·ngữ cảnh 200k·mức chi phí 5x
  • Mô hình

    Claude Opus 4.5 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#21Viết, Văn học & Ngôn ngữ#27Giải trí, Thể thao & Truyền thông#27Pháp lý & Chính phủ#34

    Batch-processing variant of Claude Opus 4.5, Anthropic's refined Opus flagship for demanding coding and reasoning work. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Nov 24, 2025·ngữ cảnh 200k·mức chi phí 2.5x
  • Mô hình

    GPT-5.1

    OpenAI🇺🇸
    Phân tích tài liệu#42Hiểu hình ảnh#49Toán học#65Pháp lý & Chính phủ#66

    GPT-5.1 is OpenAI's late-2025 successor to GPT-5, improving conversational warmth, instruction following, and reasoning quality. It suits general-purpose chat, coding, and agentic workflows that need more natural, steerable responses.

    Nov 13, 2025·ngữ cảnh 400k·mức chi phí 2x
  • Mô hình

    GPT-5.1 (batch)

    OpenAI🇺🇸
    Phân tích tài liệu#42Hiểu hình ảnh#49Toán học#65Pháp lý & Chính phủ#66

    GPT-5.1 is OpenAI's late-2025 successor to GPT-5, improving conversational warmth, instruction following, and reasoning quality for chat, coding, and agentic workflows. This is the batch variant for cheaper asynchronous processing.

    Nov 13, 2025·ngữ cảnh 400k·mức chi phí 0.9x
  • Mô hình

    GPT-5.1-Codex

    OpenAI🇺🇸

    GPT-5.1-Codex is a coding-focused variant of GPT-5.1 optimized for long-running, agentic software engineering tasks in Codex or Codex-like harnesses. It is well suited to multi-step refactors, debugging, and repository-scale coding work.

    Nov 13, 2025·ngữ cảnh 400k·mức chi phí 2x
  • Mô hình

    GPT-5.1-Codex-Mini

    OpenAI🇺🇸

    GPT-5.1-Codex-Mini is a smaller, cheaper version of GPT-5.1-Codex for agentic coding tasks. It suits lighter coding workflows where speed and cost matter more than the largest context or maximum autonomy.

    Nov 13, 2025·ngữ cảnh 400k·mức chi phí 0.4x
  • Mô hình

    Kimi K2 Thinking

    Moonshot🇨🇳

    Kimi K2 Thinking is a reasoning-focused variant of Moonshot AI's Kimi K2 that produces extended chain-of-thought before answering. It is designed for complex multi-step reasoning, math, and agentic planning tasks that benefit from deliberate thinking.

    Nov 6, 2025·ngữ cảnh 262k·mức chi phí 0.5x
  • Mô hình

    Nova Premier 1.0

    AWS🇺🇸

    Amazon Nova Premier is the most capable model in AWS's original Nova lineup, built for complex reasoning, multimodal understanding, and use as a teacher model for distillation. It targets demanding enterprise workloads requiring the highest accuracy in the Nova family.

    Oct 31, 2025·ngữ cảnh 1M·mức chi phí 2.5x
  • Mô hình

    Sonar Pro Search

    Perplexity🇺🇸

    Sonar Pro Search is a variant of Perplexity's Sonar Pro tuned for heavier, broader web search and retrieval, pulling in more sources per query. It suits research-heavy tasks where thorough web coverage matters more than raw response speed.

    Oct 30, 2025·ngữ cảnh 200k·mức chi phí 3x
  • Mô hình

    Voxtral Small 24B 2507

    Mistral🇫🇷

    Voxtral Small is Mistral AI's 24B audio-native model for speech understanding, transcription, and voice-driven instruction-following. It is designed for tasks combining spoken audio input with text reasoning, such as voice assistants and audio Q&A.

    Oct 30, 2025·ngữ cảnh 33k·mức chi phí 0.1x
  • Mô hình

    gpt-oss-safeguard-20b

    OpenAI🇺🇸

    GPT-OSS-Safeguard-20B is an open-weight variant of GPT-OSS-20B fine-tuned for safety classification and content moderation, letting developers apply and customize policy-based filtering against custom safety guidelines.

    Oct 29, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    MiniMax M2

    Minimax🇨🇳
    Toán học#196Y học & Chăm sóc sức khỏe#215Pháp lý & Chính phủ#220Phần mềm & Dịch vụ CNTT#221

    MiniMax-M2 is MiniMax's model built for coding and agentic workflows, offering strong tool-use, multi-step task execution, and software engineering performance. It is designed for developers building autonomous coding agents and complex task automation.

    Oct 23, 2025·ngữ cảnh 205k·mức chi phí 0.3x
  • Mô hình

    Qwen3 VL 32B Instruct

    Qwen🇨🇳

    Qwen3 VL 32B Instruct is a dense multimodal model in Alibaba's Qwen3 VL family, handling image understanding, OCR, and visual reasoning alongside general chat and coding ability.

    Oct 23, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Granite 4.0 Micro

    IBM🇺🇸

    Granite 4.0 H-Micro is IBM's small, efficient dense language model using a hybrid architecture, optimized for low-latency, cost-efficient workloads. It significantly outperforms earlier same-size Granite models, making it suitable for lightweight chat and instruction-following tasks.

    Oct 20, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Claude Haiku 4.5

    Anthropic🇺🇸
    Phân tích tài liệu#37Toán học#114Phần mềm & Dịch vụ CNTT#116Viết, Văn học & Ngôn ngữ#125

    Claude Haiku 4.5 is Anthropic's fast, affordable small model in the Claude 4 generation, balancing strong instruction-following and coding ability with low latency. It is designed for high-throughput agentic and chat applications where speed and cost efficiency are priorities.

    Oct 15, 2025·ngữ cảnh 200k·mức chi phí 1x
  • Mô hình

    Claude Haiku 4.5 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#37Toán học#114Phần mềm & Dịch vụ CNTT#116Viết, Văn học & Ngôn ngữ#125

    Batch-processing variant of Claude Haiku 4.5, Anthropic's fast and affordable Claude 4 small model. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Oct 15, 2025·ngữ cảnh 200k·mức chi phí 0.5x
  • Mô hình

    Qwen3 VL 8B Thinking

    Qwen🇨🇳

    The extended-reasoning variant of Qwen3 VL 8B, adding step-by-step chain-of-thought to the small multimodal model for improved accuracy on visual reasoning tasks while staying lightweight.

    Oct 14, 2025·ngữ cảnh 131k·mức chi phí 0.4x
  • Mô hình

    Qwen3 VL 8B Instruct

    Qwen🇨🇳

    Qwen3 VL 8B Instruct is a small, fast multimodal model from Alibaba's Qwen3 VL line, suited for lightweight image understanding and vision-plus-text tasks where speed and cost matter most.

    Oct 14, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Qwen3 VL 30B A3B Thinking

    Qwen🇨🇳

    The extended-reasoning variant of Qwen3 VL 30B-A3B, providing step-by-step reasoning over visual and text inputs for more complex multimodal analysis while remaining cheaper than the largest Qwen3 VL models.

    Oct 6, 2025·ngữ cảnh 262k·mức chi phí 0.4x
  • Mô hình

    Qwen3 VL 30B A3B Instruct

    Qwen🇨🇳

    Qwen3 VL 30B-A3B Instruct is a smaller mixture-of-experts multimodal model from Alibaba's Qwen3 VL family, offering solid image and document understanding at lower inference cost than the 235B version.

    Oct 6, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GPT-5 Pro

    OpenAI🇺🇸

    GPT-5 Pro is an extended-reasoning tier of GPT-5 that trades latency and cost for higher accuracy on the hardest problems. It is aimed at difficult research, coding, and analytical tasks that benefit from maximum reasoning effort.

    Oct 6, 2025·ngữ cảnh 400k·mức chi phí 23x
  • Mô hình

    GPT-5 Pro (batch)

    OpenAI🇺🇸

    GPT-5 Pro is an extended-reasoning tier of GPT-5 for the hardest research, coding, and analytical tasks. This is the batch variant for cheaper asynchronous processing.

    Oct 6, 2025·ngữ cảnh 400k·mức chi phí 12x
  • Mô hình

    GLM 4.6

    Z.AI🇨🇳
    Toán học#107Y học & Chăm sóc sức khỏe#111Giải trí, Thể thao & Truyền thông#117Pháp lý & Chính phủ#117

    GLM-4.6 is Z.ai's open-source flagship model, an incremental upgrade over GLM-4.5 with improved reasoning, coding, and agentic capabilities. It continues the GLM-4.5 architecture lineage while pushing higher benchmark performance for complex, production-grade tasks.

    Sep 30, 2025·ngữ cảnh 205k·mức chi phí 0.4x
  • Mô hình

    Claude Sonnet 4.5

    Anthropic🇺🇸
    Phân tích tài liệu#28Giải trí, Thể thao & Truyền thông#44Viết, Văn học & Ngôn ngữ#46Kinh doanh, Quản lý & Tài chính#59

    Claude Sonnet 4.5 is an upgraded mid-tier model in the Claude 4 line, improving coding, computer use, and long-context reasoning over Sonnet 4 while keeping Sonnet-level pricing. It is well suited for enterprise workflows, agentic coding, and general-purpose assistant tasks.

    Sep 29, 2025·ngữ cảnh 1M·mức chi phí 3x
  • Mô hình

    Claude Sonnet 4.5 (batch)

    Anthropic🇺🇸
    Phân tích tài liệu#28Giải trí, Thể thao & Truyền thông#44Viết, Văn học & Ngôn ngữ#46Kinh doanh, Quản lý & Tài chính#59

    Batch-processing variant of Claude Sonnet 4.5, Anthropic's upgraded mid-tier Claude 4 model. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Sep 29, 2025·ngữ cảnh 1M·mức chi phí 1.5x
  • Mô hình

    DeepSeek V3.2 Exp

    DeepSeek🇨🇳
    Y học & Chăm sóc sức khỏe#70Giải trí, Thể thao & Truyền thông#91Viết, Văn học & Ngôn ngữ#105Khoa học Sự sống, Vật lý & Xã hội#111

    DeepSeek V3.2-exp is an experimental preview checkpoint of DeepSeek V3.2, used to test architecture or training improvements ahead of the stable V3.2 release. It targets the same general reasoning, coding, and chat use cases as V3.2 with less production-hardening.

    Sep 29, 2025·ngữ cảnh 164k·mức chi phí 0.1x
  • Cydonia 24B V4.1

    thedrummer

    Cydonia 24B v4.1 is a community fine-tune from TheDrummer built on a Mistral Small 24B base, tuned for creative writing and roleplay. It favors expressive, character-driven prose over the more clinical tone of general-purpose assistant models.

    Sep 27, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Relace Apply 3

    Relace🇺🇸

    Relace Apply 3 is a specialized code-patching model that merges AI-suggested edits directly into existing source files at very high speed (thousands of tokens per second). It is designed to be paired with a larger generation model as a fast, cheap 'instant apply' layer, reducing latency and token cost in AI coding tools.

    Sep 26, 2025·ngữ cảnh 256k·mức chi phí 0.4x
  • Mô hình

    Qwen3 VL 235B A22B Thinking

    Qwen🇨🇳
    Hiểu hình ảnh#77Kinh doanh, Quản lý & Tài chính#102Pháp lý & Chính phủ#120Khoa học Sự sống, Vật lý & Xã hội#121

    The extended-reasoning variant of Qwen3 VL 235B-A22B, Alibaba's large multimodal model, which reasons step by step over visual and textual inputs for complex visual reasoning and analysis tasks.

    Sep 23, 2025·ngữ cảnh 131k·mức chi phí 0.7x
  • Mô hình

    Qwen3 VL 235B A22B Instruct

    Qwen🇨🇳
    Hiểu hình ảnh#77Kinh doanh, Quản lý & Tài chính#102Pháp lý & Chính phủ#120Khoa học Sự sống, Vật lý & Xã hội#121

    Qwen3 VL 235B-A22B Instruct is Alibaba's large-scale multimodal MoE model, combining Qwen3's language capability with vision understanding for image, document, and video tasks. It suits demanding multimodal applications needing near-flagship quality.

    Sep 23, 2025·ngữ cảnh 262k·mức chi phí 0.4x
  • Mô hình

    Qwen3 Max

    Qwen🇨🇳

    Qwen3 Max is Alibaba's flagship proprietary Qwen3 model, aimed at top-tier reasoning, coding, and agentic performance to compete with other frontier models like GPT and Claude. It is the highest-capability option in the Qwen3 lineup.

    Sep 23, 2025·ngữ cảnh 262k·mức chi phí 0.8x
  • Mô hình

    Qwen3 Coder Plus

    Qwen🇨🇳

    Qwen3 Coder Plus is a higher-capability tier of Alibaba's Qwen3 Coder family, designed for more demanding agentic coding and software engineering tasks that need stronger reasoning over large codebases than the flash/base variants.

    Sep 23, 2025·ngữ cảnh 1M·mức chi phí 0.7x
  • Mô hình

    DeepSeek V3.1 Terminus

    DeepSeek🇨🇳
    Pháp lý & Chính phủ#69Y học & Chăm sóc sức khỏe#88Viết, Văn học & Ngôn ngữ#103Khoa học Sự sống, Vật lý & Xã hội#104

    DeepSeek V3.1 Terminus is a refined checkpoint of DeepSeek V3.1, tuned for improved stability, agentic tool use, and coding performance. It is positioned as a polished general-purpose and agentic workhorse ahead of the V3.2/V4 generations.

    Sep 22, 2025·ngữ cảnh 164k·mức chi phí 0.2x
  • Mô hình

    Qwen3 Coder Flash

    Qwen🇨🇳

    Qwen3 Coder Flash is a faster, lower-cost variant of Alibaba's Qwen3 Coder model, trading some capability for speed. It suits latency-sensitive coding assistants and high-volume code-completion workloads.

    Sep 17, 2025·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    Qwen3 Next 80B A3B Thinking

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#129Kinh doanh, Quản lý & Tài chính#140Phần mềm & Dịch vụ CNTT#151Toán học#154

    The extended-reasoning variant of Qwen3-Next 80B-A3B, Alibaba's efficient next-generation MoE model, producing explicit chain-of-thought for harder reasoning and agentic tasks while retaining the architecture's low inference cost.

    Sep 11, 2025·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    Qwen3 Next 80B A3B Instruct

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#129Kinh doanh, Quản lý & Tài chính#140Phần mềm & Dịch vụ CNTT#151Toán học#154

    Qwen3-Next 80B-A3B Instruct is Alibaba's next-generation efficient mixture-of-experts model (80B total, 3B active), using a highly sparse architecture for strong general-purpose performance at low inference cost.

    Sep 11, 2025·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    Qwen Plus 0728

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#196Pháp lý & Chính phủ#198Khoa học Sự sống, Vật lý & Xã hội#200Phần mềm & Dịch vụ CNTT#222

    A dated snapshot of Alibaba's Qwen Plus mid-tier model from July 2025, offering the same general-purpose chat, reasoning, and agentic capabilities as Qwen Plus with a pinned version for reproducibility.

    Sep 8, 2025·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    Kimi K2 0905

    Moonshot🇨🇳

    Kimi K2 0905 is a September 2025 update to Moonshot AI's Kimi K2, refining agentic coding and tool-use performance. It keeps the same open-weight, trillion-parameter mixture-of-experts design tuned for long-horizon agent tasks.

    Sep 4, 2025·ngữ cảnh 262k·mức chi phí 0.5x
  • Mô hình

    Qwen3 30B A3B Thinking 2507

    Qwen🇨🇳
    Kinh doanh, Quản lý & Tài chính#164Phần mềm & Dịch vụ CNTT#167Toán học#171Y học & Chăm sóc sức khỏe#180

    The extended-reasoning (thinking) variant of the July 2025 Qwen3 30B-A3B release, producing explicit chain-of-thought for better performance on reasoning and multi-step tasks while remaining cheap to run.

    Aug 28, 2025·ngữ cảnh 82k·mức chi phí 0.4x
  • Mô hình

    Hermes 4 405B

    Nous Research🇺🇸

    Hermes 4 405B is Nous Research's next-generation fine-tune built on a Llama 3.1 405B base, improving on Hermes 3 with better reasoning, coding, and structured output while keeping a highly steerable, less-restricted persona. It suits users who want a large open-weight model with strong general capability and flexible alignment.

    Aug 26, 2025·ngữ cảnh 131k·mức chi phí 0.7x
  • Mô hình

    DeepSeek V3.1

    DeepSeek🇨🇳
    Toán học#99Viết, Văn học & Ngôn ngữ#111Y học & Chăm sóc sức khỏe#112Giải trí, Thể thao & Truyền thông#120

    DeepSeek V3.1 is an updated release of DeepSeek's V3 chat model, improving reasoning, coding, and instruction-following over earlier V3 checkpoints. It remains a low-cost, high-throughput option for general-purpose tasks.

    Aug 21, 2025·ngữ cảnh 164k·mức chi phí 0.2x
  • Mô hình

    Mistral Medium 3.1

    Mistral🇫🇷

    Mistral Medium 3.1 is a frontier-class multimodal model from Mistral AI positioned between smaller open models and top proprietary LLMs, supporting long-context reasoning and multi-image input. It has since been superseded by Mistral Medium 3.5, though it remains suitable for general reasoning and vision tasks.

    Aug 13, 2025·ngữ cảnh 131k·mức chi phí 0.4x
  • Mô hình

    Mistral Medium 3.1 (batch)

    Mistral🇫🇷

    This is the batch-processing variant of Mistral Medium 3.1 (mistral-medium-3.1), Mistral AI's proprietary enterprise-grade model (131K-token context, multimodal text, image, and PDF input) released in August 2025. It delivers frontier-level coding, STEM reasoning, and enterprise performance at roughly 8x lower cost than top-tier models, and the batch endpoint further cuts cost via asynchronous processing for high-volume, non-latency-sensitive use.

    Aug 13, 2025·ngữ cảnh 131k·mức chi phí 0.2x
  • Mô hình

    GLM 4.5V

    Z.AI🇨🇳
    Hiểu hình ảnh#110Y học & Chăm sóc sức khỏe#195Pháp lý & Chính phủ#208Phần mềm & Dịch vụ CNTT#209

    GLM-4.5V is Z.ai's vision-language model built on the GLM-4.5-Air architecture (106B total, 12B active), designed for versatile multimodal reasoning. It handles complex STEM problem-solving, GUI agent tasks, and video understanding via a 3D convolutional vision encoder.

    Aug 11, 2025·ngữ cảnh 66k·mức chi phí 0.4x
  • Mô hình

    GPT-5

    OpenAI🇺🇸
    Hiểu hình ảnh#70Pháp lý & Chính phủ#77Y học & Chăm sóc sức khỏe#95Toán học#100

    GPT-5 is OpenAI's flagship model released in 2025, unifying fast responses and deep reasoning in a single system. It is built for advanced coding, agentic workflows, and complex multi-step problem solving with strong instruction following.

    Aug 7, 2025·ngữ cảnh 400k·mức chi phí 2x
  • Mô hình

    GPT-5 (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#70Pháp lý & Chính phủ#77Y học & Chăm sóc sức khỏe#95Toán học#100

    GPT-5 is OpenAI's flagship model unifying fast responses and deep reasoning, built for advanced coding, agentic workflows, and complex problem solving. This is the batch variant for cheaper asynchronous processing.

    Aug 7, 2025·ngữ cảnh 400k·mức chi phí 0.9x
  • Mô hình

    GPT-5 Mini

    OpenAI🇺🇸
    Hiểu hình ảnh#97Toán học#161Y học & Chăm sóc sức khỏe#176Giải trí, Thể thao & Truyền thông#179

    GPT-5 Mini is a smaller, faster, and cheaper version of GPT-5 from OpenAI. It targets high-volume chat, lightweight coding, and agentic tasks where speed and cost matter more than maximum reasoning depth.

    Aug 7, 2025·ngữ cảnh 400k·mức chi phí 0.4x
  • Mô hình

    GPT-5 Mini (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#97Toán học#161Y học & Chăm sóc sức khỏe#176Giải trí, Thể thao & Truyền thông#179

    GPT-5 Mini is a smaller, faster, and cheaper version of GPT-5 from OpenAI, suited to high-volume chat, lightweight coding, and agentic tasks. This is the batch variant for cheaper asynchronous processing.

    Aug 7, 2025·ngữ cảnh 400k·mức chi phí 0.2x
  • Mô hình

    GPT-5 Nano

    OpenAI🇺🇸
    Hiểu hình ảnh#112Toán học#210Y học & Chăm sóc sức khỏe#217Kinh doanh, Quản lý & Tài chính#229

    GPT-5 Nano is the smallest and cheapest model in the GPT-5 family, optimized for extremely fast, low-latency responses. It suits simple classification, extraction, and high-throughput tasks where cost and speed dominate.

    Aug 7, 2025·ngữ cảnh 400k·mức chi phí 0.1x
  • Mô hình

    GPT-5 Nano (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#112Toán học#210Y học & Chăm sóc sức khỏe#217Kinh doanh, Quản lý & Tài chính#229

    GPT-5 Nano is the smallest and cheapest model in the GPT-5 family, optimized for fast, low-latency, high-throughput tasks. This is the batch variant for cheaper asynchronous processing.

    Aug 7, 2025·ngữ cảnh 400k·mức chi phí 0.1x
  • Mô hình

    gpt-oss-120b

    OpenAI🇺🇸
    Toán học#183Kinh doanh, Quản lý & Tài chính#208Y học & Chăm sóc sức khỏe#213Phần mềm & Dịch vụ CNTT#216

    GPT-OSS-120B is OpenAI's open-weight mixture-of-experts model (around 120B parameters) released under a permissive license for self-hosted deployment. It offers strong reasoning and coding performance for teams that need to run OpenAI-quality models on their own infrastructure.

    Aug 5, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    gpt-oss-120b (batch)

    OpenAI🇺🇸
    Toán học#183Kinh doanh, Quản lý & Tài chính#208Y học & Chăm sóc sức khỏe#213Phần mềm & Dịch vụ CNTT#216

    This is the batch-processing variant of gpt-oss-120b, OpenAI's open-weight mixture-of-experts model (117B total parameters, 5.1B active, Apache 2.0 license, 128K-token context, released August 5, 2025), offering the same configurable reasoning depth and native tool-use design at a lower cost via OpenRouter's asynchronous batch endpoint. It fits high-volume, non-latency-sensitive reasoning and agentic workloads where throughput economics matter more than response time.

    Aug 5, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    gpt-oss-20b

    OpenAI🇺🇸
    Toán học#226Y học & Chăm sóc sức khỏe#229Kinh doanh, Quản lý & Tài chính#239Phần mềm & Dịch vụ CNTT#242

    GPT-OSS-20B is OpenAI's smaller open-weight model (around 20B parameters), designed to run on more modest hardware while retaining solid reasoning and coding ability for local or self-hosted use cases.

    Aug 5, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    gpt-oss-20b (batch)

    OpenAI🇺🇸
    Toán học#226Y học & Chăm sóc sức khỏe#229Kinh doanh, Quản lý & Tài chính#239Phần mềm & Dịch vụ CNTT#242

    This is the batch-processing variant of gpt-oss-20b, OpenAI's open-weight mixture-of-experts model (21B total parameters, 3.6B active, Apache 2.0 license, 128K-token context, released August 5, 2025), offering the same configurable reasoning effort at a lower cost via OpenRouter's asynchronous batch endpoint. It suits high-volume, non-latency-sensitive workloads like bulk summarization or classification rather than interactive tool-calling use.

    Aug 5, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Claude Opus 4.1

    Anthropic🇺🇸
    Giải trí, Thể thao & Truyền thông#61Viết, Văn học & Ngôn ngữ#65Y học & Chăm sóc sức khỏe#73Pháp lý & Chính phủ#80

    Claude Opus 4.1 is an incremental upgrade to Claude Opus 4, improving coding accuracy, reasoning, and agentic task performance. It remains Anthropic's top-tier model for complex, high-stakes engineering and analysis work in that generation.

    Aug 5, 2025·ngữ cảnh 200k·mức chi phí 15x
  • Mô hình

    Claude Opus 4.1 (batch)

    Anthropic🇺🇸
    Giải trí, Thể thao & Truyền thông#61Viết, Văn học & Ngôn ngữ#65Y học & Chăm sóc sức khỏe#73Pháp lý & Chính phủ#80

    Batch-processing variant of Claude Opus 4.1, Anthropic's top-tier model for complex coding and reasoning. Same capabilities as the standard model, offered at lower cost via asynchronous batch processing.

    Aug 5, 2025·ngữ cảnh 200k·mức chi phí 7.5x
  • Mô hình

    Codestral 2508

    Mistral🇫🇷

    Codestral is Mistral AI's dedicated coding model, trained on 80+ programming languages for code generation, completion, and refactoring. It is optimized for low-latency code assistant workflows including fill-in-the-middle completion.

    Aug 1, 2025·ngữ cảnh 256k·mức chi phí 0.2x
  • Mô hình

    Codestral 2508 (batch)

    Mistral🇫🇷

    This is the asynchronous batch-serving variant of Mistral's Codestral 2508, a dense code-focused model (256K-token context, released end of July 2025 under Mistral's Non-Production License) specialized for fill-in-the-middle completion, code correction, and test generation. The batch endpoint processes requests asynchronously at a discount versus the standard endpoint, fitting large, non-interactive code-generation jobs rather than latency-sensitive use.

    Aug 1, 2025·ngữ cảnh 256k·mức chi phí 0.1x
  • Mô hình

    Qwen3 Coder 30B A3B Instruct

    Qwen🇨🇳

    A 30B-A3B mixture-of-experts version of Qwen3 Coder, Alibaba's agentic coding model, offering strong code generation and repository-scale understanding at lower inference cost than the full-size Coder model.

    Jul 31, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Qwen3 30B A3B Instruct 2507

    Qwen🇨🇳
    Kinh doanh, Quản lý & Tài chính#164Phần mềm & Dịch vụ CNTT#167Toán học#171Y học & Chăm sóc sức khỏe#180

    A July 2025 instruction-tuned refresh of Qwen3 30B-A3B, Alibaba's efficient MoE model, with improved instruction following for general chat and task completion at low cost.

    Jul 29, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    GLM 4.5

    Z.AI🇨🇳
    Y học & Chăm sóc sức khỏe#133Giải trí, Thể thao & Truyền thông#138Toán học#138Kinh doanh, Quản lý & Tài chính#139

    GLM-4.5 is Z.ai's (Zhipu AI) open-source flagship MoE model (355B total, 32B active) designed for complex reasoning, coding, and agentic workflows. It serves as the base architecture for the broader GLM-4.5 family, including the smaller Air and vision-enabled V variants.

    Jul 25, 2025·ngữ cảnh 131k·mức chi phí 0.5x
  • Mô hình

    GLM 4.5 Air

    Z.AI🇨🇳
    Toán học#177Khoa học Sự sống, Vật lý & Xã hội#185Phần mềm & Dịch vụ CNTT#186Kinh doanh, Quản lý & Tài chính#187

    GLM-4.5-Air is a lighter-weight, more efficient variant of Z.ai's GLM-4.5 model, trading some capacity for lower cost and faster inference. It is well suited for general reasoning and coding tasks where a smaller footprint is preferred over the full GLM-4.5 model.

    Jul 25, 2025·ngữ cảnh 131k·mức chi phí 0.2x
  • Mô hình

    Qwen3 235B A22B Thinking 2507

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#100Kinh doanh, Quản lý & Tài chính#112Phần mềm & Dịch vụ CNTT#117Khoa học Sự sống, Vật lý & Xã hội#120

    The extended-reasoning variant of the July 2025 Qwen3 235B-A22B release, which generates explicit chain-of-thought for harder reasoning, math, and multi-step agentic tasks before producing a final answer.

    Jul 25, 2025·ngữ cảnh 131k·mức chi phí 0.4x
  • Mô hình

    Qwen3 Coder 480B A35B

    Qwen🇨🇳
    Phần mềm & Dịch vụ CNTT#158Pháp lý & Chính phủ#166Kinh doanh, Quản lý & Tài chính#169Giải trí, Thể thao & Truyền thông#170

    Qwen3 Coder is Alibaba's agentic coding model built on the Qwen3 architecture, tuned for large-context, repository-level code understanding, generation, and tool use. It targets coding agents and IDE assistants that need to reason across whole codebases.

    Jul 23, 2025·ngữ cảnh 262k·mức chi phí 0.2x
  • Mô hình

    UI-TARS 7B

    Bytedance🇨🇳

    UI-TARS-1.5-7B is ByteDance Seed's multimodal vision-language agent model built for GUI interaction across desktop, web, and mobile interfaces. It achieves strong results on computer-use benchmarks like OSWorld and AndroidWorld, making it suited for automated screen-based agent tasks rather than general chat.

    Jul 22, 2025·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    Gemini 2.5 Flash Lite

    Google🇺🇸

    Gemini 2.5 Flash-Lite is Google's most lightweight and cost-effective model in the 2.5 family, optimized for very high-throughput, low-latency tasks such as classification, extraction, and simple chat. It trades some reasoning depth for speed and low cost.

    Jul 22, 2025·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Gemini 2.5 Flash Lite (batch)

    Google🇺🇸

    Gemini 2.5 Flash-Lite is Google's most lightweight and cost-effective model in the 2.5 family, optimized for very high-throughput, low-latency tasks such as classification, extraction, and simple chat. This is the batch variant for asynchronous, lower-cost processing.

    Jul 22, 2025·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Qwen3 235B A22B Instruct 2507

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#100Kinh doanh, Quản lý & Tài chính#112Phần mềm & Dịch vụ CNTT#117Khoa học Sự sống, Vật lý & Xã hội#120

    A July 2025 refreshed release of Qwen3 235B-A22B, Alibaba's large MoE model, with improved instruction following and general capability over the original version. Suited for demanding general-purpose, coding, and agentic workloads.

    Jul 21, 2025·ngữ cảnh 262k·mức chi phí 0.1x
  • Mô hình

    Kimi K2 0711

    Moonshot🇨🇳

    Kimi K2 is Moonshot AI's open-weight trillion-parameter mixture-of-experts model built for agentic tasks, coding, and tool use. It is designed as a cost-efficient alternative to top proprietary models for long-horizon agent workflows.

    Jul 11, 2025·ngữ cảnh 131k·mức chi phí 0.5x
  • Uncensored

    cognitivecomputations

    Dolphin Mistral 24B Venice Edition is an uncensored community fine-tune of Mistral 24B, created by Cognitive Computations in collaboration with Venice.ai. It targets general conversation, coding assistance, and creative writing without the guardrails of typical instruction-tuned models.

    Jul 9, 2025·ngữ cảnh 128k·mức chi phí 0.2x
  • Mô hình

    Hunyuan A13B Instruct

    Tencent🇨🇳

    Hunyuan-A13B-Instruct is Tencent's open-source fine-grained MoE model (80B total, 13B active) that adapts its reasoning depth in real time, using a fast path for simple queries and deeper multistep reasoning for complex ones. It is tuned for math, science, and agent-based tasks with a 256k context window.

    Jul 8, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Morph V3 Large

    Morph🇺🇸

    Morph V3 Large is Morph's high-accuracy code-merging "apply" model, achieving around 98% accuracy on complex edits at moderate speed. It is best used when merge correctness matters more than raw speed, such as applying intricate multi-file diffs from a coding agent.

    Jul 7, 2025·ngữ cảnh 262k·mức chi phí 0.5x
  • Mô hình

    Morph V3 Fast

    Morph🇺🇸

    Morph V3 Fast is a specialized "apply" model from Morph that merges AI-generated code edits into source files at high speed with around 96% accuracy. It is designed to pair with a coding agent to quickly and reliably apply proposed code diffs rather than generate code itself.

    Jul 7, 2025·ngữ cảnh 82k·mức chi phí 0.3x
  • Mô hình

    ERNIE 4.5 VL 424B A47B

    Baidu🇨🇳

    ERNIE 4.5 VL 424B A47B is Baidu's multimodal mixture-of-experts model, with 424B total and 47B active parameters per token, jointly trained on text and image data. It supports both thinking and non-thinking modes for vision-language reasoning, image understanding, and long-context generation up to 131K tokens in English and Chinese.

    Jun 30, 2025·ngữ cảnh 123k·mức chi phí 0.3x
  • Mô hình

    Mistral Small 3.2 24B

    Mistral🇫🇷

    Mistral Small 3.2 is an incremental update to Mistral Small 3.1 from Mistral AI, improving instruction-following and reducing repetition errors. It remains a 24B multimodal model suited for efficient chat, coding, and vision tasks.

    Jun 20, 2025·ngữ cảnh 256k·mức chi phí 0.1x
  • Mô hình

    MiniMax M1

    Minimax🇨🇳
    Toán học#186Y học & Chăm sóc sức khỏe#192Khoa học Sự sống, Vật lý & Xã hội#199Giải trí, Thể thao & Truyền thông#200

    MiniMax-M1 is MiniMax's large language model built for reasoning, coding, and agentic tool use with efficient long-context processing. It targets developers needing strong reasoning performance at competitive cost.

    Jun 17, 2025·ngữ cảnh 1M·mức chi phí 0.5x
  • Mô hình

    Gemini 2.5 Flash

    Google🇺🇸
    Hiểu hình ảnh#76Viết, Văn học & Ngôn ngữ#123Pháp lý & Chính phủ#132Giải trí, Thể thao & Truyền thông#133

    Gemini 2.5 Flash is Google's fast, cost-efficient multimodal model built for high-volume everyday tasks like chat, summarization, and lightweight agentic workflows. It balances speed and quality with configurable thinking for tasks that need a bit more reasoning.

    Jun 17, 2025·ngữ cảnh 1M·mức chi phí 0.5x
  • Mô hình

    Gemini 2.5 Flash (batch)

    Google🇺🇸
    Hiểu hình ảnh#76Viết, Văn học & Ngôn ngữ#123Pháp lý & Chính phủ#132Giải trí, Thể thao & Truyền thông#133

    Gemini 2.5 Flash is Google's fast, cost-efficient multimodal model built for high-volume everyday tasks like chat, summarization, and lightweight agentic workflows. This is the batch variant for asynchronous, lower-cost processing.

    Jun 17, 2025·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    Gemini 2.5 Pro

    Google🇺🇸
    Phân tích tài liệu#36Hiểu hình ảnh#52Pháp lý & Chính phủ#57Viết, Văn học & Ngôn ngữ#63

    Gemini 2.5 Pro is Google's flagship reasoning model of the 2.5 generation, designed for complex multimodal understanding, advanced coding, and long-context analysis (up to 1M tokens). It excels at multi-step problem solving and agentic tasks that require deep reasoning.

    Jun 17, 2025·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    Gemini 2.5 Pro (batch)

    Google🇺🇸
    Phân tích tài liệu#36Hiểu hình ảnh#52Pháp lý & Chính phủ#57Viết, Văn học & Ngôn ngữ#63

    Gemini 2.5 Pro is Google's flagship reasoning model of the 2.5 generation, designed for complex multimodal understanding, advanced coding, and long-context analysis. This is the batch variant for asynchronous, lower-cost processing.

    Jun 17, 2025·ngữ cảnh 1M·mức chi phí 0.9x
  • Mô hình

    o3 Pro

    OpenAI🇺🇸

    o3-pro is a higher-compute version of OpenAI's o3 reasoning model, applying extra thinking effort for maximum reliability on the hardest math, science, and coding problems.

    Jun 10, 2025·ngữ cảnh 200k·mức chi phí 17x
  • Mô hình

    Gemini 2.5 Pro Preview 06-05

    Google🇺🇸

    Gemini 2.5 Pro Preview is the preview release of Google's flagship 2.5 Pro reasoning model, designed for complex multimodal understanding, advanced coding, and long-context analysis. It offers the same strong multi-step reasoning and agentic capabilities as the GA release.

    Jun 5, 2025·ngữ cảnh 1M·mức chi phí 2x
  • Mô hình

    R1 0528

    DeepSeek🇨🇳
    Y học & Chăm sóc sức khỏe#123Giải trí, Thể thao & Truyền thông#125Phần mềm & Dịch vụ CNTT#126Toán học#142

    DeepSeek R1 (0528) is an updated checkpoint of DeepSeek's R1 reasoning model with improved chain-of-thought quality and benchmark performance. It targets the same complex math, coding, and logic reasoning use cases as the original R1.

    May 28, 2025·ngữ cảnh 164k·mức chi phí 0.4x
  • Mô hình

    Claude Sonnet 4

    Anthropic🇺🇸

    Claude Sonnet 4 is Anthropic's balanced mid-tier model from the Claude 4 generation, offering strong coding, reasoning, and agentic performance at lower cost than Opus. It is designed as a versatile default for everyday software engineering and assistant tasks.

    May 22, 2025·ngữ cảnh 200k·mức chi phí 3x
  • Mô hình

    Mistral Medium 3

    Mistral🇫🇷

    Mistral Medium 3 is a mid-sized frontier model from Mistral AI balancing strong reasoning, coding, and multimodal performance against lower cost than Mistral Large. It targets enterprise and agentic use cases needing near-flagship quality at reduced price.

    May 7, 2025·ngữ cảnh 131k·mức chi phí 0.4x
  • Mô hình

    Llama Guard 4 12B

    Meta🇺🇸

    Llama Guard 4 12B is Meta's open-weight safety classifier model used to moderate and filter both prompts and model outputs across text and image content. It is designed to be deployed alongside other Llama models to detect unsafe or policy-violating content rather than for general chat.

    Apr 30, 2025·ngữ cảnh 164k·mức chi phí 0.1x
  • Mô hình

    Qwen3 30B A3B

    Qwen🇨🇳
    Kinh doanh, Quản lý & Tài chính#164Phần mềm & Dịch vụ CNTT#167Toán học#171Y học & Chăm sóc sức khỏe#180

    Qwen3 30B-A3B is a smaller mixture-of-experts model (30B total, 3B active parameters) in Alibaba's Qwen3 line, offering good general-purpose capability at low inference cost thanks to its sparse activation.

    Apr 28, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Qwen3 8B

    Qwen🇨🇳

    Qwen3 8B is a small, fast open-weight model in Alibaba's Qwen3 family, suited for lightweight chat, simple coding, and general tasks where speed and low cost matter more than top-end capability.

    Apr 28, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Qwen3 14B

    Qwen🇨🇳

    Qwen3 14B is a mid-size open-weight model in Alibaba's Qwen3 generation, supporting hybrid reasoning and non-reasoning modes. It is designed for general-purpose chat, coding, and agentic tasks at a moderate compute cost.

    Apr 28, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Qwen3 32B

    Qwen🇨🇳
    Toán học#127Khoa học Sự sống, Vật lý & Xã hội#203Y học & Chăm sóc sức khỏe#203Phần mềm & Dịch vụ CNTT#204

    Qwen3 32B is a dense open-weight model from Alibaba's Qwen3 generation, supporting both reasoning and non-reasoning modes. It is a solid general-purpose choice for chat, coding, and agentic tasks where a mid-size dense model is preferred over MoE.

    Apr 28, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Qwen3 235B A22B

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#100Kinh doanh, Quản lý & Tài chính#112Phần mềm & Dịch vụ CNTT#117Khoa học Sự sống, Vật lý & Xã hội#120

    Qwen3 235B-A22B is a large mixture-of-experts model (235B total, 22B active parameters) from Alibaba's Qwen3 family, offering strong multilingual reasoning, coding, and agentic capabilities. Its MoE design gives near-flagship quality at lower inference cost than an equivalent dense model.

    Apr 28, 2025·ngữ cảnh 131k·mức chi phí 0.4x
  • Mô hình

    o4 Mini High

    OpenAI🇺🇸
    Hiểu hình ảnh#83Toán học#146Y học & Chăm sóc sức khỏe#163Pháp lý & Chính phủ#171

    o4-mini-high is OpenAI's o4-mini reasoning model run at a higher reasoning-effort setting, improving accuracy on math, coding, and visual reasoning tasks at the cost of some speed.

    Apr 16, 2025·ngữ cảnh 200k·mức chi phí 0.9x
  • Mô hình

    o3

    OpenAI🇺🇸
    Y học & Chăm sóc sức khỏe#65Hiểu hình ảnh#75Toán học#88Pháp lý & Chính phủ#92

    OpenAI o3 is a next-generation reasoning model succeeding o1, with stronger performance on math, science, coding, and multi-step logical problem solving through deeper internal deliberation.

    Apr 16, 2025·ngữ cảnh 200k·mức chi phí 1.5x
  • Mô hình

    o3 (batch)

    OpenAI🇺🇸
    Y học & Chăm sóc sức khỏe#65Hiểu hình ảnh#75Toán học#88Pháp lý & Chính phủ#92

    OpenAI o3 is a reasoning model with strong math, science, coding, and multi-step logical problem-solving performance. This is the batch variant for cheaper asynchronous processing.

    Apr 16, 2025·ngữ cảnh 200k·mức chi phí 0.8x
  • Mô hình

    o4 Mini

    OpenAI🇺🇸
    Hiểu hình ảnh#83Toán học#146Y học & Chăm sóc sức khỏe#163Pháp lý & Chính phủ#171

    OpenAI o4-mini is an efficient reasoning model succeeding o3-mini, combining strong math, coding, and visual reasoning performance with low cost and fast response times.

    Apr 16, 2025·ngữ cảnh 200k·mức chi phí 0.9x
  • Mô hình

    o4 Mini (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#83Toán học#146Y học & Chăm sóc sức khỏe#163Pháp lý & Chính phủ#171

    OpenAI o4-mini is an efficient reasoning model with strong math, coding, and visual reasoning performance at low cost. This is the batch variant for cheaper asynchronous processing.

    Apr 16, 2025·ngữ cảnh 200k·mức chi phí 0.5x
  • Mô hình

    GPT-4.1

    OpenAI🇺🇸
    Hiểu hình ảnh#78Giải trí, Thể thao & Truyền thông#108Viết, Văn học & Ngôn ngữ#115Pháp lý & Chính phủ#116

    GPT-4.1 is OpenAI's 2025 flagship model, improving coding, instruction-following, and long-context handling (up to 1M tokens) over GPT-4o. It is well suited for agentic coding, complex reasoning, and tasks requiring very long context.

    Apr 14, 2025·ngữ cảnh 1M·mức chi phí 1.5x
  • Mô hình

    GPT-4.1 (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#78Giải trí, Thể thao & Truyền thông#108Viết, Văn học & Ngôn ngữ#115Pháp lý & Chính phủ#116

    GPT-4.1 is OpenAI's 2025 flagship model, improving coding, instruction-following, and long-context handling (up to 1M tokens) over GPT-4o, well suited for agentic coding and complex reasoning. This is the batch variant for asynchronous, lower-cost processing.

    Apr 14, 2025·ngữ cảnh 1M·mức chi phí 0.8x
  • Mô hình

    GPT-4.1 Mini

    OpenAI🇺🇸
    Hiểu hình ảnh#84Pháp lý & Chính phủ#175Giải trí, Thể thao & Truyền thông#180Viết, Văn học & Ngôn ngữ#181

    GPT-4.1 Mini is a smaller, faster, cheaper variant of GPT-4.1 that retains strong coding and reasoning ability with a large context window. It is suited for high-volume or latency-sensitive applications that still need solid general capability.

    Apr 14, 2025·ngữ cảnh 1M·mức chi phí 0.3x
  • Mô hình

    GPT-4.1 Mini (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#84Pháp lý & Chính phủ#175Giải trí, Thể thao & Truyền thông#180Viết, Văn học & Ngôn ngữ#181

    GPT-4.1 Mini is a smaller, faster, cheaper variant of GPT-4.1 that retains strong coding and reasoning ability with a large context window, suited for high-volume or latency-sensitive applications. This is the batch variant for asynchronous, lower-cost processing.

    Apr 14, 2025·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    GPT-4.1 Nano

    OpenAI🇺🇸
    Hiểu hình ảnh#131Pháp lý & Chính phủ#223Phần mềm & Dịch vụ CNTT#238Khoa học Sự sống, Vật lý & Xã hội#238

    GPT-4.1 Nano is OpenAI's smallest and fastest GPT-4.1 variant, optimized for extremely low-latency, low-cost tasks such as classification, extraction, and simple chat. It sacrifices some capability for speed and price relative to Mini and full GPT-4.1.

    Apr 14, 2025·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    GPT-4.1 Nano (batch)

    OpenAI🇺🇸
    Hiểu hình ảnh#131Pháp lý & Chính phủ#223Phần mềm & Dịch vụ CNTT#238Khoa học Sự sống, Vật lý & Xã hội#238

    GPT-4.1 Nano is OpenAI's smallest and fastest GPT-4.1 variant, optimized for extremely low-latency, low-cost tasks such as classification, extraction, and simple chat. This is the batch variant for asynchronous, lower-cost processing.

    Apr 14, 2025·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Llama 4 Maverick

    Meta🇺🇸
    Hiểu hình ảnh#113Y học & Chăm sóc sức khỏe#235Viết, Văn học & Ngôn ngữ#236Giải trí, Thể thao & Truyền thông#236

    Llama 4 Maverick is Meta's high-capability open-weight Mixture-of-Experts model, offering strong multimodal (text and image) understanding, reasoning, and coding at efficient inference cost. It targets general-purpose assistant and agentic use cases needing frontier-adjacent quality in an open model.

    Apr 5, 2025·ngữ cảnh 1M·mức chi phí 0.1x
  • Mô hình

    Llama 4 Scout

    Meta🇺🇸
    Hiểu hình ảnh#119Pháp lý & Chính phủ#237Toán học#242Giải trí, Thể thao & Truyền thông#243

    Llama 4 Scout is Meta's efficient open-weight Mixture-of-Experts model with an industry-leading long context window, designed for tasks requiring extensive context handling alongside multimodal understanding. It offers a lighter, faster alternative to Llama 4 Maverick.

    Apr 5, 2025·ngữ cảnh 1.3M·mức chi phí 0.1x
  • Mô hình

    DeepSeek V3 0324

    DeepSeek🇨🇳
    Giải trí, Thể thao & Truyền thông#137Viết, Văn học & Ngôn ngữ#148Pháp lý & Chính phủ#152Y học & Chăm sóc sức khỏe#154

    DeepSeek V3 (0324) is a checkpoint of DeepSeek's V3 mixture-of-experts model, offering strong general reasoning, coding, and instruction-following. It is a cost-efficient open-weight alternative to closed frontier models.

    Mar 24, 2025·ngữ cảnh 164k·mức chi phí 0.2x
  • Mô hình

    o1-pro

    OpenAI🇺🇸

    o1-pro is a higher-compute version of OpenAI's o1 reasoning model, applying more thinking effort for greater reliability on the hardest math, science, and coding problems.

    Mar 19, 2025·ngữ cảnh 200k·mức chi phí 125x
  • Mô hình

    Mistral Small 3.1 24B

    Mistral🇫🇷
    Hiểu hình ảnh#120Giải trí, Thể thao & Truyền thông#263Phần mềm & Dịch vụ CNTT#266Toán học#267

    Mistral Small 3.1 is a 24B open-weight multimodal model from Mistral AI supporting vision input alongside text, with a large context window. It is built for fast, cost-effective chat, coding, and image-understanding tasks.

    Mar 17, 2025·ngữ cảnh 128k·mức chi phí 0.2x
  • Mô hình

    Gemma 3 4B

    Google🇺🇸
    Y học & Chăm sóc sức khỏe#226Kinh doanh, Quản lý & Tài chính#233Pháp lý & Chính phủ#242Viết, Văn học & Ngôn ngữ#263

    Gemma 3 4B IT is a small open-weight instruction-tuned model from Google's Gemma 3 family, designed for fast, low-cost inference on edge or resource-constrained hardware. It handles basic chat, summarization, and simple multimodal tasks efficiently.

    Mar 13, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Gemma 3 12B

    Google🇺🇸
    Kinh doanh, Quản lý & Tài chính#177Pháp lý & Chính phủ#192Khoa học Sự sống, Vật lý & Xã hội#211Toán học#217

    Gemma 3 12B IT is a mid-size open-weight instruction-tuned model from Google's Gemma 3 family, supporting multimodal (text and image) input and long context. It is well suited for general-purpose chat, reasoning, and coding assistance on self-hosted or resource-constrained deployments.

    Mar 13, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Command A

    Cohere🇨🇦
    Viết, Văn học & Ngôn ngữ#193Giải trí, Thể thao & Truyền thông#195Pháp lý & Chính phủ#200Y học & Chăm sóc sức khỏe#207

    Command A is Cohere's flagship enterprise model, optimized for agentic tool use, RAG, multilingual business tasks, and instruction-following at efficient serving cost. It is designed for enterprise deployments needing strong accuracy with lower compute requirements than typical frontier models.

    Mar 13, 2025·ngữ cảnh 256k·mức chi phí 2x
  • Mô hình

    Reka Flash 3

    Reka AI🇺🇸

    Reka Flash 3 is a mid-size multimodal model from Reka AI balancing speed and capability, supporting text and vision inputs for general chat, reasoning, and multimodal tasks at a favorable cost-performance ratio.

    Mar 12, 2025·ngữ cảnh 66k·mức chi phí 0.1x
  • Mô hình

    Gemma 3 27B

    Google🇺🇸
    Hiểu hình ảnh#105Y học & Chăm sóc sức khỏe#178Kinh doanh, Quản lý & Tài chính#185Viết, Văn học & Ngôn ngữ#187

    Gemma 3 27B IT is the largest dense model in Google's open-weight Gemma 3 family, supporting multimodal input and strong multilingual capability. It targets general chat, reasoning, and coding tasks where higher quality than smaller Gemma variants is needed while staying open and self-hostable.

    Mar 12, 2025·ngữ cảnh 131k·mức chi phí 0.1x
  • Skyfall 36B V2

    thedrummer

    Skyfall 36B v2 is a larger creative-writing and roleplay fine-tune from TheDrummer, offering more capacity and nuance than the smaller Rocinante and Cydonia models in the same lineup. It is designed for expressive, character-consistent storytelling.

    Mar 10, 2025·ngữ cảnh 33k·mức chi phí 0.2x
  • Mô hình

    Sonar Reasoning Pro

    Perplexity🇺🇸

    Sonar Reasoning Pro combines Perplexity's web-grounded search with explicit chain-of-thought reasoning. It is designed for complex analytical questions that require both up-to-date factual grounding and multi-step logical reasoning over the retrieved information.

    Mar 7, 2025·ngữ cảnh 128k·mức chi phí 1.5x
  • Mô hình

    Sonar Pro

    Perplexity🇺🇸

    Sonar Pro is Perplexity's higher-quality, search-grounded model with a larger context window and more thorough citation handling than the base Sonar. It is designed for complex, multi-part queries that need accurate, well-sourced, up-to-date answers.

    Mar 7, 2025·ngữ cảnh 200k·mức chi phí 3x
  • Mô hình

    Sonar Deep Research

    Perplexity🇺🇸

    Sonar Deep Research is Perplexity's autonomous research model that performs multi-step web searches, reads many sources, and synthesizes long, well-cited reports. It is suited for comprehensive research tasks, literature-style summaries, and in-depth fact-finding that goes far beyond a single search query.

    Mar 7, 2025·ngữ cảnh 128k·mức chi phí 1.5x
  • Mô hình

    Saba

    Mistral🇫🇷

    Mistral Saba is a regional model from Mistral AI tailored for Arabic and languages of the Middle East and South Asia. It is designed for culturally and linguistically accurate text generation and chat in those regions.

    Feb 17, 2025·ngữ cảnh 33k·mức chi phí 0.1x
  • Mô hình

    o3 Mini High

    OpenAI🇺🇸
    Toán học#151Phần mềm & Dịch vụ CNTT#189Viết, Văn học & Ngôn ngữ#203Khoa học Sự sống, Vật lý & Xã hội#207

    o3-mini-high is OpenAI's o3-mini reasoning model run at a higher reasoning-effort setting, trading some speed for improved accuracy on math, coding, and logic tasks.

    Feb 12, 2025·ngữ cảnh 200k·mức chi phí 0.9x
  • Mô hình

    Aion-RP 1.0 (8B)

    Aion Labs🇵🇱

    Aion-RP-Llama-3.1-8B is a lightweight roleplay-focused fine-tune of Meta's Llama 3.1 8B from AionLabs. It is designed for fast, low-cost character roleplay and creative fiction rather than complex reasoning or coding.

    Feb 4, 2025·ngữ cảnh 33k·mức chi phí 0.4x
  • Mô hình

    Qwen2.5 VL 72B Instruct

    Qwen🇨🇳
    Hiểu hình ảnh#123

    Qwen 2.5 VL 72B Instruct is Alibaba's large open-weight vision-language model, capable of image understanding, document/OCR parsing, and visual reasoning alongside text. It is designed for multimodal tasks that combine strong language ability with detailed visual comprehension.

    Feb 1, 2025·ngữ cảnh 128k·mức chi phí 0.3x
  • Mô hình

    Qwen-Plus

    Qwen🇨🇳
    Y học & Chăm sóc sức khỏe#196Pháp lý & Chính phủ#198Khoa học Sự sống, Vật lý & Xã hội#200Phần mềm & Dịch vụ CNTT#222

    Qwen Plus is Alibaba's proprietary mid-tier model in the Qwen lineup, balancing cost and capability for general chat, reasoning, and agentic tasks. It sits between the lightweight and flagship Qwen offerings for everyday production use.

    Feb 1, 2025·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    o3 Mini

    OpenAI🇺🇸
    Toán học#151Phần mềm & Dịch vụ CNTT#189Viết, Văn học & Ngôn ngữ#203Khoa học Sự sống, Vật lý & Xã hội#207

    o3-mini is a smaller, faster, and cheaper version of OpenAI's o3 reasoning model, aimed at math, coding, and logic tasks where lower latency and cost matter more than maximum reasoning depth.

    Jan 31, 2025·ngữ cảnh 200k·mức chi phí 0.9x
  • Mô hình

    o3 Mini (batch)

    OpenAI🇺🇸
    Toán học#151Phần mềm & Dịch vụ CNTT#189Viết, Văn học & Ngôn ngữ#203Khoa học Sự sống, Vật lý & Xã hội#207

    o3-mini is a smaller, faster, and cheaper version of OpenAI's o3 reasoning model for math, coding, and logic tasks. This is the batch variant for cheaper asynchronous processing.

    Jan 31, 2025·ngữ cảnh 200k·mức chi phí 0.5x
  • Mô hình

    Mistral Small 3

    Mistral🇫🇷
    Toán học#285Y học & Chăm sóc sức khỏe#289Phần mềm & Dịch vụ CNTT#297Khoa học Sự sống, Vật lý & Xã hội#303

    Mistral Small 3 (24B, January 2025) is a dense open-weight model from Mistral AI tuned for low-latency chat, coding, and instruction-following. It is designed to deliver strong performance at a fraction of the cost of larger models.

    Jan 30, 2025·ngữ cảnh 33k·mức chi phí 0.1x
  • Mô hình

    Sonar

    Perplexity🇺🇸

    Sonar is Perplexity's fast, search-augmented language model that grounds answers in real-time web results with citations. It is designed for quick, factual question answering and up-to-date information retrieval rather than deep reasoning or creative tasks.

    Jan 27, 2025·ngữ cảnh 127k·mức chi phí 0.3x
  • Mô hình

    R1

    DeepSeek🇨🇳
    Y học & Chăm sóc sức khỏe#123Giải trí, Thể thao & Truyền thông#125Phần mềm & Dịch vụ CNTT#126Toán học#142

    DeepSeek R1 is DeepSeek's dedicated reasoning model, trained with reinforcement learning to produce explicit chain-of-thought before answering. It is designed for complex math, coding, and logic problems, and was notable for matching top closed reasoning models at a fraction of the cost.

    Jan 20, 2025·ngữ cảnh 64k·mức chi phí 0.5x
  • Mô hình

    MiniMax-01

    Minimax🇨🇳

    MiniMax-01 is MiniMax's foundational large language model, an early entry in the MiniMax series offering general-purpose chat, reasoning, and long-context handling. It established the base architecture later refined into the MiniMax M-series.

    Jan 15, 2025·ngữ cảnh 1M·mức chi phí 0.2x
  • Mô hình

    Phi 4

    Microsoft🇺🇸
    Toán học#275Pháp lý & Chính phủ#294Y học & Chăm sóc sức khỏe#299Phần mềm & Dịch vụ CNTT#307

    Phi-4 is Microsoft's small open-weight language model that punches above its parameter count, particularly strong on reasoning and math tasks relative to its size. It is well suited for efficient, cost-conscious deployments needing solid reasoning without a large model footprint.

    Jan 10, 2025·ngữ cảnh 16k·mức chi phí 0.1x
  • Mô hình

    DeepSeek V3

    DeepSeek🇨🇳
    Giải trí, Thể thao & Truyền thông#137Viết, Văn học & Ngôn ngữ#148Pháp lý & Chính phủ#152Y học & Chăm sóc sức khỏe#154

    DeepSeek Chat is DeepSeek's general-purpose conversational model, corresponding to the non-thinking mode of the DeepSeek V3 model family. It is designed for everyday chat, coding assistance, and reasoning tasks at low cost.

    Dec 26, 2024·ngữ cảnh 164k·mức chi phí 0.2x
  • Llama 3.3 Euryale 70B

    sao10k

    L3.3 Euryale 70B is a Llama 3.3 70B fine-tune by Sao10K, continuing the Euryale line's focus on immersive roleplay and creative writing with improved base-model capability.

    Dec 18, 2024·ngữ cảnh 131k·mức chi phí 0.2x
  • Mô hình

    o1

    OpenAI🇺🇸
    Hiểu hình ảnh#90Giải trí, Thể thao & Truyền thông#122Viết, Văn học & Ngôn ngữ#129Toán học#144

    OpenAI o1 is the first model in OpenAI's dedicated reasoning line, using extended internal chain-of-thought before answering. It is designed for hard math, science, and coding problems that benefit from deliberate step-by-step reasoning.

    Dec 17, 2024·ngữ cảnh 200k·mức chi phí 13x
  • Mô hình

    Command R7B (12-2024)

    Cohere🇨🇦

    Command R7B (12-2024) is Cohere's compact model in the Command R family, designed for fast, low-cost RAG, tool use, and conversational tasks on resource-constrained deployments. It trades some capability for speed and efficiency versus the larger Command R models.

    Dec 14, 2024·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    Llama 3.3 70B Instruct

    Meta🇺🇸
    Pháp lý & Chính phủ#235Y học & Chăm sóc sức khỏe#240Khoa học Sự sống, Vật lý & Xã hội#249Giải trí, Thể thao & Truyền thông#253

    Llama 3.3 70B Instruct is Meta's open-weight model that delivers performance close to Llama 3.1 405B at a fraction of the size and cost, with a 128K context window. It is well suited for general chat, reasoning, and coding in self-hosted deployments.

    Dec 6, 2024·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Nova Lite 1.0

    AWS🇺🇸

    Amazon Nova Lite is a low-cost multimodal model from AWS's original Nova family, handling text, image, and video input with fast response times. It is designed for high-volume, latency-sensitive tasks like summarization and simple assistant workflows.

    Dec 5, 2024·ngữ cảnh 300k·mức chi phí 0.1x
  • Mô hình

    Nova Micro 1.0

    AWS🇺🇸

    Amazon Nova Micro is a text-only, ultra-low-latency model from AWS's Nova family, optimized purely for speed and cost efficiency. It suits simple, high-throughput text tasks like classification, extraction, and short-form generation.

    Dec 5, 2024·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    Nova Pro 1.0

    AWS🇺🇸

    Amazon Nova Pro is a balanced multimodal model from AWS's Nova family, handling text, image, and video input with strong accuracy at moderate cost and latency. It is designed as an all-round workhorse for agentic and enterprise applications.

    Dec 5, 2024·ngữ cảnh 300k·mức chi phí 0.7x
  • Mô hình

    GPT-4o (2024-11-20)

    OpenAI🇺🇸

    GPT-4o 2024-11-20 is a pinned November 2024 snapshot of OpenAI's GPT-4o with improved creative writing quality and instruction-following. It retains the same fast multimodal capabilities as GPT-4o for general-purpose chat and vision tasks.

    Nov 20, 2024·ngữ cảnh 128k·mức chi phí 2x
  • Mô hình

    Mistral Large 2407

    Mistral🇫🇷
    Giải trí, Thể thao & Truyền thông#240Pháp lý & Chính phủ#245Viết, Văn học & Ngôn ngữ#252Toán học#263

    Mistral Large 2 (2407) is Mistral AI's mid-2024 flagship model with strong multilingual, coding, and reasoning capabilities and a 128K context window. It was built as a high-capability alternative to top proprietary models for enterprise use cases.

    Nov 19, 2024·ngữ cảnh 131k·mức chi phí 1x
  • Mô hình

    Qwen2.5 Coder 32B Instruct

    Qwen🇨🇳
    Pháp lý & Chính phủ#284Toán học#287Kinh doanh, Quản lý & Tài chính#289Phần mềm & Dịch vụ CNTT#290

    Qwen 2.5 Coder 32B Instruct is Alibaba's dedicated code-generation model, specialized for writing, completing, and explaining code across many programming languages. It is well suited for coding assistants and developer tooling that need strong code-specific performance.

    Nov 11, 2024·ngữ cảnh 33k·mức chi phí 0.3x
  • UnslopNemo 12B

    thedrummer

    UnslopNemo 12B is a TheDrummer fine-tune of Mistral Nemo 12B aimed at reducing repetitive, generic "AI-isms" (slop) in generated prose. It is built for creative writing and roleplay where more natural, varied language is desired.

    Nov 8, 2024·ngữ cảnh 1M·mức chi phí 0.1x
  • Magnum v4 72B

    anthracite-org

    Magnum v4 72B is a community fine-tune from Anthracite based on Qwen2.5 72B, trained to emulate the prose quality and creative writing style of Claude 3 models. It is aimed at long-form creative writing and roleplay rather than reasoning or coding tasks.

    Oct 22, 2024·ngữ cảnh 33k·mức chi phí 1x
  • Mô hình

    Qwen2.5 7B Instruct

    Qwen🇨🇳

    Qwen 2.5 7B Instruct is a compact open-weight chat model from Alibaba's Qwen family, tuned for fast, low-cost general-purpose tasks. It trades some capability for speed and efficiency compared to the larger Qwen 2.5 models.

    Oct 16, 2024·ngữ cảnh 33k·mức chi phí 0.1x
  • Mô hình

    Llama 3.2 1B Instruct

    Meta🇺🇸
    Toán học#357Pháp lý & Chính phủ#370Y học & Chăm sóc sức khỏe#376Giải trí, Thể thao & Truyền thông#390

    Llama 3.2 1B Instruct is Meta's smallest open-weight instruction-tuned model, built for extremely low-latency, on-device or edge deployment. It handles basic text tasks like classification and simple chat rather than complex reasoning.

    Sep 25, 2024·ngữ cảnh 60k·mức chi phí 0.1x
  • Mô hình

    Llama 3.2 3B Instruct

    Meta🇺🇸
    Y học & Chăm sóc sức khỏe#339Toán học#342Giải trí, Thể thao & Truyền thông#354Pháp lý & Chính phủ#355

    Llama 3.2 3B Instruct is a small open-weight instruction-tuned model from Meta, suited for on-device and edge deployment where efficiency matters. It handles lightweight chat, summarization, and simple instruction-following tasks.

    Sep 25, 2024·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Qwen2.5 72B Instruct

    Qwen🇨🇳
    Toán học#248Pháp lý & Chính phủ#253Phần mềm & Dịch vụ CNTT#268Kinh doanh, Quản lý & Tài chính#271

    Qwen 2.5 72B Instruct is Alibaba's large general-purpose open-weight chat model, strong at multilingual conversation, reasoning, and coding. It suits demanding general assistant tasks where a bigger dense model improves quality over smaller Qwen variants.

    Sep 19, 2024·ngữ cảnh 33k·mức chi phí 0.1x
  • Mô hình

    Command R (08-2024)

    Cohere🇨🇦
    Y học & Chăm sóc sức khỏe#287Pháp lý & Chính phủ#292Khoa học Sự sống, Vật lý & Xã hội#305Toán học#309

    Command R (08-2024) is a mid-sized model from Cohere's Command R family, optimized for retrieval-augmented generation, tool use, and multilingual conversational tasks. It balances capability and cost for enterprise RAG and chat applications.

    Aug 30, 2024·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    Command R+ (08-2024)

    Cohere🇨🇦
    Pháp lý & Chính phủ#277Y học & Chăm sóc sức khỏe#277Giải trí, Thể thao & Truyền thông#284Viết, Văn học & Ngôn ngữ#285

    Command R+ (08-2024) is Cohere's larger, more capable model in the Command R family, built for complex RAG pipelines, multi-step tool use, and enterprise-grade multilingual tasks. It offers higher accuracy than Command R at increased cost.

    Aug 30, 2024·ngữ cảnh 128k·mức chi phí 2x
  • Llama 3.1 Euryale 70B v2.2

    sao10k

    L3.1 Euryale 70B is a Llama 3.1 70B fine-tune by Sao10K, optimized for roleplay, creative writing, and character-driven storytelling with rich, expressive prose.

    Aug 28, 2024·ngữ cảnh 131k·mức chi phí 0.3x
  • Mô hình

    Hermes 3 70B Instruct

    Nous Research🇺🇸

    Hermes 3 70B is Nous Research's fine-tune of Meta's Llama 3.1 70B, offering the same steerable, low-refusal instruction-following and roleplay focus as the 405B version in a smaller, faster package. It suits general chat, creative writing, and agentic use where lower cost matters.

    Aug 18, 2024·ngữ cảnh 131k·mức chi phí 0.2x
  • Mô hình

    Hermes 3 405B Instruct

    Nous Research🇺🇸

    Hermes 3 405B is Nous Research's fine-tune of Meta's Llama 3.1 405B, tuned for strong instruction-following, roleplay, and agentic reasoning with reduced refusals. It suits users who want a large, highly steerable open-weight model with fewer alignment restrictions than the base model.

    Aug 16, 2024·ngữ cảnh 131k·mức chi phí 0.3x
  • Llama 3 8B Lunaris

    sao10k

    L3 Lunaris 8B is a community fine-tune of Meta's Llama 3 8B by Sao10K, tuned for roleplay and creative writing. It is designed for immersive character-driven storytelling rather than general assistant tasks.

    Aug 13, 2024·ngữ cảnh 8k·mức chi phí 0.1x
  • Mô hình

    GPT-4o (2024-08-06)

    OpenAI🇺🇸
    Hiểu hình ảnh#126Giải trí, Thể thao & Truyền thông#202Viết, Văn học & Ngôn ngữ#204Pháp lý & Chính phủ#227

    GPT-4o 2024-08-06 is a pinned August 2024 snapshot of OpenAI's GPT-4o that added support for structured JSON outputs. It retains GPT-4o's fast multimodal text, image, and audio capabilities for general-purpose chat and vision tasks.

    Aug 6, 2024·ngữ cảnh 128k·mức chi phí 2x
  • Mô hình

    Llama 3.1 70B Instruct

    Meta🇺🇸
    Pháp lý & Chính phủ#258Y học & Chăm sóc sức khỏe#273Khoa học Sự sống, Vật lý & Xã hội#277Toán học#278

    Llama 3.1 70B Instruct is Meta's open-weight instruction-tuned model, offering strong general reasoning, chat, and coding performance with a 128K context window. It is a popular, well-balanced choice for self-hosted deployments needing near-frontier quality at open-weight cost.

    Jul 23, 2024·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Llama 3.1 8B Instruct

    Meta🇺🇸
    Y học & Chăm sóc sức khỏe#325Pháp lý & Chính phủ#327Phần mềm & Dịch vụ CNTT#334Kinh doanh, Quản lý & Tài chính#334

    Llama 3.1 8B Instruct is Meta's small open-weight instruction-tuned model, designed for fast, low-cost chat and general text tasks with a 128K context window. It suits lightweight applications and resource-constrained self-hosted deployments.

    Jul 23, 2024·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    Mistral Nemo

    Mistral🇫🇷

    Mistral Nemo is a 12B-parameter multilingual model built by Mistral AI in collaboration with NVIDIA, featuring a 128K context window. It is designed for efficient reasoning, coding, and multilingual chat with a strong performance-to-size ratio.

    Jul 19, 2024·ngữ cảnh 131k·mức chi phí 0.1x
  • Mô hình

    GPT-4o-mini

    OpenAI🇺🇸

    GPT-4o Mini is OpenAI's small, low-cost multimodal model, offering solid reasoning and vision capability at a fraction of GPT-4o's price. It is well suited for high-volume, cost-sensitive applications like chatbots and lightweight assistants.

    Jul 18, 2024·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    GPT-4o-mini (2024-07-18)

    OpenAI🇺🇸
    Hiểu hình ảnh#129Pháp lý & Chính phủ#230Viết, Văn học & Ngôn ngữ#233Y học & Chăm sóc sức khỏe#254

    GPT-4o Mini 2024-07-18 is the pinned initial-release snapshot of OpenAI's small, low-cost GPT-4o Mini model from July 2024. It offers the same fast, affordable multimodal reasoning and vision capability as GPT-4o Mini, kept fixed for reproducibility.

    Jul 18, 2024·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    GPT-4o-mini (batch)

    OpenAI🇺🇸

    GPT-4o Mini is OpenAI's small, low-cost multimodal model, offering solid reasoning and vision capability at a fraction of GPT-4o's price, well suited for high-volume, cost-sensitive applications. This is the batch variant for asynchronous, lower-cost processing.

    Jul 18, 2024·ngữ cảnh 128k·mức chi phí 0.1x
  • Mô hình

    Gemma 2 27B

    Google🇺🇸
    Viết, Văn học & Ngôn ngữ#262Giải trí, Thể thao & Truyền thông#266Pháp lý & Chính phủ#283Y học & Chăm sóc sức khỏe#285

    Gemma 2 27B IT is Google's open-weight instruction-tuned language model from the Gemma 2 family, suitable for general chat, reasoning, and text generation tasks that can be self-hosted. It offers strong quality for its size while remaining runnable on modest hardware.

    Jul 13, 2024·ngữ cảnh 8k·mức chi phí 0.2x
  • Mô hình

    GPT-4o

    OpenAI🇺🇸

    GPT-4o is OpenAI's natively multimodal flagship model handling text, image, and audio input with fast response times. It is designed for general-purpose chat, vision tasks, and real-time conversational applications.

    May 13, 2024·ngữ cảnh 128k·mức chi phí 2x
  • Mô hình

    GPT-4o (2024-05-13)

    OpenAI🇺🇸
    Hiểu hình ảnh#102Giải trí, Thể thao & Truyền thông#189Viết, Văn học & Ngôn ngữ#195Pháp lý & Chính phủ#205

    GPT-4o 2024-05-13 is the pinned initial-release snapshot of OpenAI's natively multimodal GPT-4o model from May 2024. It offers the same fast text, image, and audio understanding as GPT-4o, kept fixed for reproducibility.

    May 13, 2024·ngữ cảnh 128k·mức chi phí 3x
  • Mô hình

    GPT-4o (batch)

    OpenAI🇺🇸

    GPT-4o is OpenAI's original natively multimodal flagship model, handling text, vision, and audio in one system. It is a solid general-purpose choice for chat, vision tasks, and everyday coding; this is the batch variant for cheaper asynchronous processing.

    May 13, 2024·ngữ cảnh 128k·mức chi phí 1x
  • Mô hình

    Mixtral 8x22B Instruct

    Mistral🇫🇷

    Mixtral 8x22B is Mistral AI's sparse mixture-of-experts model with 8 experts of 22B parameters each and roughly 39B active parameters. It offers strong reasoning, multilingual, and coding performance at efficient inference cost relative to a dense model of similar quality.

    Apr 17, 2024·ngữ cảnh 66k·mức chi phí 1x
  • Mô hình

    WizardLM-2 8x22B

    Microsoft🇺🇸

    WizardLM-2 8x22B is Microsoft's open-weight Mixture-of-Experts model tuned for complex instruction following, reasoning, and multilingual chat. It offers strong general-purpose capability among open models, suited for advanced conversational and reasoning tasks.

    Apr 16, 2024·ngữ cảnh 66k·mức chi phí 0.2x
  • Mô hình

    GPT-4 Turbo

    OpenAI🇺🇸

    GPT-4 Turbo is OpenAI's faster, cheaper evolution of GPT-4 with a larger context window and a more recent knowledge cutoff. It is suited for complex reasoning, coding, and long-document tasks at lower cost than the original GPT-4.

    Apr 9, 2024·ngữ cảnh 128k·mức chi phí 6.5x
  • Mô hình

    GPT-4 Turbo (batch)

    OpenAI🇺🇸

    GPT-4 Turbo is OpenAI's faster, cheaper evolution of GPT-4 with a larger context window, suited for complex reasoning, coding, and long-document tasks. This is the batch variant for asynchronous, lower-cost processing.

    Apr 9, 2024·ngữ cảnh 128k·mức chi phí 3x
  • Mô hình

    Mistral Large

    Mistral🇫🇷
    Giải trí, Thể thao & Truyền thông#240Pháp lý & Chính phủ#245Viết, Văn học & Ngôn ngữ#252Toán học#263

    Mistral Large is Mistral AI's flagship general-purpose model line for complex reasoning, coding, and multilingual tasks, with this key generally pointing at the current default Mistral Large release. It is designed for demanding enterprise workloads requiring strong reasoning and instruction-following.

    Feb 26, 2024·ngữ cảnh 128k·mức chi phí 1x
  • Mô hình

    GPT-3.5 Turbo (older v0613)

    OpenAI🇺🇸
    Viết, Văn học & Ngôn ngữ#319Giải trí, Thể thao & Truyền thông#322Toán học#326Pháp lý & Chính phủ#326

    GPT-3.5 Turbo 0613 is a pinned June 2023 snapshot of OpenAI's GPT-3.5 Turbo chat model, kept for reproducibility. It offers the same fast, low-cost chat capability as GPT-3.5 Turbo, suited for simple conversational and summarization tasks.

    Jan 25, 2024·ngữ cảnh 4k·mức chi phí 0.5x
  • Mô hình

    Auto Router

    OpenRouter🇺🇸

    Auto Router is OpenRouter's task-aware meta-router: it classifies each prompt into one of roughly 30 task types and routes it to the model the community spends most on for that task over a trailing 7-day window, filtered by a cost_tier dial (low by default) and kept sticky across a conversation's turns. Routing is free and each response is priced at the underlying routed model's rate, making it a good fit for mixed workloads where the best model varies from request to request.

    Nov 8, 2023·ngữ cảnh 2M·mức chi phí -
  • Mô hình

    GPT-3.5 Turbo Instruct

    OpenAI🇺🇸
    Viết, Văn học & Ngôn ngữ#319Giải trí, Thể thao & Truyền thông#322Toán học#326Pháp lý & Chính phủ#326

    GPT-3.5 Turbo Instruct is OpenAI's completion-style (non-chat) variant of GPT-3.5 Turbo, designed for legacy instruction-following completion APIs rather than multi-turn chat. It is suited for simple text completion and instruction tasks in legacy integrations.

    Sep 28, 2023·ngữ cảnh 4k·mức chi phí 0.6x
  • Mô hình

    GPT-3.5 Turbo 16k

    OpenAI🇺🇸

    GPT-3.5 Turbo 16K is a variant of OpenAI's GPT-3.5 Turbo with an extended 16K-token context window, otherwise sharing the same fast, low-cost chat capabilities. It suits tasks needing longer input or output than the standard 4K context allows.

    Aug 28, 2023·ngữ cảnh 16k·mức chi phí 1x
  • Weaver (alpha)

    mancer

    Weaver (alpha) from Mancer is a fine-tune on top of Qwen2.5 72B designed to recreate Claude-style prose verbosity for roleplay and narrative writing. It is aimed at creative and roleplay use cases rather than coherence-critical or agentic tasks, with a modest 8K context window.

    Aug 2, 2023·ngữ cảnh 8k·mức chi phí 0.2x
  • ReMM SLERP 13B

    undi95

    RemM SLERP L2 13B is a community model merge (using SLERP interpolation) built on Llama 2 13B by well-known model merger Undi95. It is a classic roleplay- and creative-writing-oriented model rather than a general assistant.

    Jul 22, 2023·ngữ cảnh 6k·mức chi phí 0.2x
  • MythoMax 13B

    gryphe

    MythoMax L2 13B is a popular community fine-tune (Gryphe) built on Llama 2 13B, blending several roleplay and storytelling models. It is designed for creative writing, roleplay, and character-driven chat rather than technical or agentic tasks.

    Jul 2, 2023·ngữ cảnh 8k·mức chi phí 0.1x
  • Mô hình

    GPT-3.5 Turbo

    OpenAI🇺🇸
    Viết, Văn học & Ngôn ngữ#319Giải trí, Thể thao & Truyền thông#322Toán học#326Pháp lý & Chính phủ#326

    GPT-3.5 Turbo is OpenAI's earlier-generation chat model, fast and inexpensive but less capable than GPT-4-class models. It is suited for simple chat, summarization, and lightweight tasks where cost and speed outweigh the need for advanced reasoning.

    May 28, 2023·ngữ cảnh 16k·mức chi phí 0.3x
  • Mô hình

    GPT-3.5 Turbo (batch)

    OpenAI🇺🇸
    Viết, Văn học & Ngôn ngữ#319Giải trí, Thể thao & Truyền thông#322Toán học#326Pháp lý & Chính phủ#326

    GPT-3.5 Turbo is OpenAI's earlier-generation chat model, fast and inexpensive, suited for simple chat, summarization, and lightweight tasks. This is the batch variant for asynchronous, lower-cost processing.

    May 28, 2023·ngữ cảnh 16k·mức chi phí 0.2x
  • Mô hình

    GPT-4

    OpenAI🇺🇸

    GPT-4 is OpenAI's landmark large multimodal model offering strong reasoning, coding, and instruction-following well beyond GPT-3.5. It is suited for complex tasks requiring deeper reasoning, though it has largely been superseded by faster, cheaper GPT-4o and GPT-4.1 models.

    May 28, 2023·ngữ cảnh 8k·mức chi phí 15x