NVIDIA: Nemotron Nano 9B V2

nvidia/nemotron-nano-9b-v2

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so.

Model specifications

Input
text
Output
text
Context
131,072 tokens
Max output
131,072 tokens
Input price
$0.04 / 1M tokens
Output price
$0.16 / 1M tokens
Released
2026-03-19

Capabilities

  • Streaming
  • Function calling
  • Vision
  • JSON mode
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · tier input average $0.04 / 1M tokens · tier output average $0.16 / 1M tokens

DeepInfra

Tier: Standard · Region: US · Quantization: fp16

Pricing
Input
$0.04 / 1M tokens
Output
$0.16 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
US
Zero data retention
Yes
Data retention
Zero retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
Yes
HIPAA compliant
No
SOC 2 certified
Yes
BYOK supported
Yes

Frequently asked questions

What is NVIDIA: Nemotron Nano 9B V2?
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so.
How much does NVIDIA: Nemotron Nano 9B V2 cost?
Input costs start at $0.04 / 1M tokens and output costs start at $0.16 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of NVIDIA: Nemotron Nano 9B V2?
NVIDIA: Nemotron Nano 9B V2 supports a 131,072 token context window and up to 131,072 output tokens.
What capabilities does NVIDIA: Nemotron Nano 9B V2 support?
NVIDIA: Nemotron Nano 9B V2 supports Streaming, Function calling, Vision, JSON mode, Playground.
Which providers offer NVIDIA: Nemotron Nano 9B V2?
NVIDIA: Nemotron Nano 9B V2 is available from DeepInfra.
How do providers handle data privacy for NVIDIA: Nemotron Nano 9B V2?
1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models