Qwen: Qwen3 32B

qwen/qwen3-32b

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, coding, and logical inference, and a "non-thinking" mode for faster, general-purpose conversation. The model demonstrates strong performance in instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.

Model specifications

Context
40,960 tokens
Max output
32,768 tokens
Input price
$0.08 / 1M tokens
Output price
$0.24 / 1M tokens
Released
2026-02-05

Capabilities

  • Streaming
  • Playground

Provider pricing, discounts and data privacy

Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.

Standard service tier

2 available providers · tier input average $0.13 / 1M tokens · tier output average $0.92 / 1M tokens

AtlasCloud

Tier: Standard · Region: US

Pricing
Input
$0.1 / 1M tokens
Output
.2 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
US
Zero data retention
No
Data retention
7-day retention
Used for training
No
Data collection
Moderated
No
GDPR compliant
No
HIPAA compliant
No
SOC 2 certified
No
BYOK supported
No

Alibaba Cloud Int.(SG)

Tier: Standard · Region: SG · Quantization: int8

Pricing
Input
$0.16 / 1M tokens
Output
$0.64 / 1M tokens

No provider discount is currently published.

Data privacy and compliance
Region
SG
Zero data retention
No
Data retention
30-day retention
Used for training
No
Data collection
Moderated
Yes
GDPR compliant
Yes
HIPAA compliant
No
SOC 2 certified
Yes
BYOK supported
Yes

Frequently asked questions

What is Qwen: Qwen3 32B?
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, coding, and logical inference, and a "non-thinking" mode for faster, general-purpose conversation. The model demonstrates strong performance in instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.
How much does Qwen: Qwen3 32B cost?
Input costs start at $0.08 / 1M tokens and output costs start at $0.24 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of Qwen: Qwen3 32B?
Qwen: Qwen3 32B supports a 40,960 token context window and up to 32,768 output tokens.
What capabilities does Qwen: Qwen3 32B support?
Qwen: Qwen3 32B supports Streaming, Playground.
Which providers offer Qwen: Qwen3 32B?
Qwen: Qwen3 32B is available from AtlasCloud, Alibaba Cloud Int.(SG).
How do providers handle data privacy for Qwen: Qwen3 32B?
2 of 2 providers report that customer data is not used for training, and 0 offer zero-data-retention routing. Retention, compliance, and privacy-policy links are listed per provider.

Browse all AI models