nemotron-lightning-3.5-30b-a3b
by Requesty · family nemotron · listed Aug 2026RT
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
Current rates USD per 1M tokens
Input$0.045
Output$0.18
Cache read$0.009
Cache write—
Against the market
On the record
- Model id
nemotron-lightning-3.5-30b-a3b- Context window
- 262K tokens
- Max output
- 262K tokens
- Input modalities
- text
- Output modalities
- text
- Reasoning
- yes
- Tool calls
- yes
- Open weights
- no
- Knowledge cutoff
- —
- Released
- 2026-08-15
- Last updated
- 2026-08-15
- Provider docs
- requesty.ai/solution/llm-routing/models
Back of the envelope
≈