Nemotron 3.5 Lightning
NVIDIA🇺🇸
nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning is NVIDIA's open 30B mixture-of-experts model with 3B active parameters and a 1M-token context window, built for fast, accurate execution within long-running agent pipelines. It is designed as a low-latency, low-cost execution layer for always-on agentic systems.
费用比率
0.1x
上下文
262K
发布日期
2026年8月11日
输入
Text
输出
Text
支持
推理工具调用结构化输出
最擅长
该模型排名最高的类别。
未找到排名靠前的类别。
性能
根据近期请求测得的中位延迟和吞吐量。
吞吐量301 tok/s
延迟0.55s