NVIDIA: Nemotron Nano 12B 2 VL

nvidia/nemotron-nano-12b-v2-vl

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s memory-efficient sequence modeling for significantly higher throughput and lower latency. The model supports inputs of text and multi-image documents, producing natural-language outputs. It is trained on high-quality NVIDIA-curated synthetic datasets optimized for optical-character recognition, chart reasoning, and multimodal comprehension. Nemotron Nano 2 VL achieves leading results on OCRBench v2 and scores ≈ 74 average across MMMU, MathVista, AI2D, OCRBench, OCR-Reasoning, ChartQA, DocVQA, and Video-MME—surpassing prior open VL baselines. With Efficient Video Sampling (EVS), it handles long-form videos while reducing inference cost. Open-weights, training data, and fine-tuning recipes are released under a permissive NVIDIA open license, with deployment supported across NeMo, NIM, and major inference runtimes.

Model specifications

Input
text
Output
text
Context
131,072 tokens
Max output
131,072 tokens
Input price
$0.2 / 1M tokens
Output price
$0.6 / 1M tokens
Released
2026-03-19

Capabilities

  • Streaming
  • Function calling
  • Vision
  • JSON mode
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

1 available provider · tier input average $0.2 / 1M tokens · tier output average $0.6 / 1M tokens

DeepInfra

Tier: Standard · Region: US · Quantization: fp8

Pricing
Input
$0.2 / 1M tokens
Output
$0.6 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
US
Zero data retention
Yes
Data retention
Zero retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
Yes
HIPAA compliant
No
SOC 2 certified
Yes
BYOK supported
Yes

Frequently asked questions

What is NVIDIA: Nemotron Nano 12B 2 VL?
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s memory-efficient sequence modeling for significantly higher throughput and lower latency. The model supports inputs of text and multi-image documents, producing natural-language outputs. It is trained on high-quality NVIDIA-curated synthetic datasets optimized for optical-character recognition, chart reasoning, and multimodal comprehension. Nemotron Nano 2 VL achieves leading results on OCRBench v2 and scores ≈ 74 average across MMMU, MathVista, AI2D, OCRBench, OCR-Reasoning, ChartQA, DocVQA, and Video-MME—surpassing prior open VL baselines. With Efficient Video Sampling (EVS), it handles long-form videos while reducing inference cost. Open-weights, training data, and fine-tuning recipes are released under a permissive NVIDIA open license, with deployment supported across NeMo, NIM, and major inference runtimes.
How much does NVIDIA: Nemotron Nano 12B 2 VL cost?
Input costs start at $0.2 / 1M tokens and output costs start at $0.6 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of NVIDIA: Nemotron Nano 12B 2 VL?
NVIDIA: Nemotron Nano 12B 2 VL supports a 131,072 token context window and up to 131,072 output tokens.
What capabilities does NVIDIA: Nemotron Nano 12B 2 VL support?
NVIDIA: Nemotron Nano 12B 2 VL supports Streaming, Function calling, Vision, JSON mode, Playground.
Which providers offer NVIDIA: Nemotron Nano 12B 2 VL?
NVIDIA: Nemotron Nano 12B 2 VL is available from DeepInfra.
How do providers handle data privacy for NVIDIA: Nemotron Nano 12B 2 VL?
1 of 1 providers report that customer data is not used for training, and 1 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models