NVIDIA: Nemotron 3 Super

nvidia/nemotron-3-super-120b-a12b

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models. The model features a 1M token context window for long-term agent coherence, cross-document reasoning, and multi-step task planning. Latent MoE enables calling 4 experts for the inference cost of only one, improving intelligence and generalization. Multi-environment RL training across 10+ environments delivers leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified. Fully open with weights, datasets, and recipes under the NVIDIA Open License, Nemotron 3 Super allows easy customization and secure deployment anywhere — from workstation to cloud.

Model specifications

Input
text
Output
text
Context
262,144 tokens
Max output
262,144 tokens
Input price
$0.1 / 1M tokens
Output price
$0.5 / 1M tokens
Released
2026-03-19

Capabilities

  • Streaming
  • Function calling
  • Vision
  • JSON mode
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · tier input average $0.1 / 1M tokens · tier output average $0.5 / 1M tokens

DeepInfra

Tier: Standard · Region: US · Quantization: fp16

Pricing
Input
$0.1 / 1M tokens
Output
$0.5 / 1M tokens
Cache read
$0.04 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
US
Zero data retention
Yes
Data retention
Zero retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
Yes
HIPAA compliant
No
SOC 2 certified
Yes
BYOK supported
Yes

Frequently asked questions

What is NVIDIA: Nemotron 3 Super?
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer Mixture-of-Experts architecture with multi-token prediction (MTP), it delivers over 50% higher token generation compared to leading open models. The model features a 1M token context window for long-term agent coherence, cross-document reasoning, and multi-step task planning. Latent MoE enables calling 4 experts for the inference cost of only one, improving intelligence and generalization. Multi-environment RL training across 10+ environments delivers leading accuracy on benchmarks including AIME 2025, TerminalBench, and SWE-Bench Verified. Fully open with weights, datasets, and recipes under the NVIDIA Open License, Nemotron 3 Super allows easy customization and secure deployment anywhere — from workstation to cloud.
How much does NVIDIA: Nemotron 3 Super cost?
Input costs start at $0.1 / 1M tokens and output costs start at $0.5 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of NVIDIA: Nemotron 3 Super?
NVIDIA: Nemotron 3 Super supports a 262,144 token context window and up to 262,144 output tokens.
What capabilities does NVIDIA: Nemotron 3 Super support?
NVIDIA: Nemotron 3 Super supports Streaming, Function calling, Vision, JSON mode, Playground.
Which providers offer NVIDIA: Nemotron 3 Super?
NVIDIA: Nemotron 3 Super is available from DeepInfra.
How do providers handle data privacy for NVIDIA: Nemotron 3 Super?
1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models