All Models
nemotron-lightning-3.5-30b-a3b
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
Available Providers (1)
| Provider | Model ID | Input Cost | Output Cost | Context | Max Output | Docs |
|---|---|---|---|---|---|---|
| | nemotron-lightning-3.5-30b-a3b | $0.04/MTok | $0.18/MTok | 262.1K | 262.1K |
Capabilities
Reasoning
Tool Calling
Attachments
Open Weights
Structured Output