nemotron-3.5-lightning-30b-a3b
by Requesty · family nemotron · listed Aug 2026RT
NVIDIA Nemotron 3.5 Lightning 30B-A3B is a hybrid Mamba-2 + MoE + Attention model with 30B total and 3B active parameters, pre-trained on over 20T tokens with an NVFP4 recipe and Multi-Token Prediction for fast generation. Up to 1M token context for long-running autonomous agents, sub-agent workhorse deployments, and agentic workflows. Supports reasoning and tool calling. English and coding languages plus Spanish, French, German, Italian, and Japanese. Open weights under the OpenMDW License Agreement v1.1. Part of the NVIDIA Nemotron family.
Current rates USD per 1M tokens
Inputfree
Outputfree
Cache read—
Cache write—
Against the market
On the record
- Model id
nemotron-3.5-lightning-30b-a3b- Context window
- 1.0M tokens
- Max output
- 66K tokens
- Input modalities
- text
- Output modalities
- text
- Reasoning
- yes
- Tool calls
- yes
- Open weights
- no
- Knowledge cutoff
- —
- Released
- 2026-08-11
- Last updated
- 2026-08-11
- Provider docs
- requesty.ai/solution/llm-routing/models
Back of the envelope
≈
Also listed at
- Nvidiafree in · free out
- Merge Gatewayfree in · free out