Nemotron 3.5 Lightning
NVIDIA🇺🇸
nvidia/nemotron-3.5-lightningNemotron 3.5 Lightning is NVIDIA's open 30B mixture-of-experts model with 3B active parameters and a 1M-token context window, built for fast, accurate execution within long-running agent pipelines. It is designed as a low-latency, low-cost execution layer for always-on agentic systems.
Cost rate
0.1x
Context
262K
Released
Aug 11, 2026
Input
Text
Output
Text
Support
ReasoningTool callingStructured outputs
Best at
The categories where this model ranks highest.
No top-ranked categories found.
Performance
Median latency and throughput measured across recent requests.
Throughput301 tok/s
Latency