Qwen: Qwen3 235B A22B Instruct 2507
qwen/qwen3-235b-a22b-instruct-2507
Qwen3-235B-A22B-Instruct-2507 has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 235B in total and 22B activated Number of Paramaters (Non-Embedding): 234B Number of Layers: 94 Number of Attention Heads (GQA): 64 for Q and 4 for KV Number of Experts: 128 Number of Activated Experts: 8 Context Length: 262,144 natively and extendable up to 1,010,000 tokens NOTE: This model supports only non-thinking mode and does not generate <think></think> blocks in its output. Meanwhile, specifying enable_thinking=False is no longer required.
Model specifications
- Context
- 131,072 tokens
- Max output
- 8,192 tokens
- Input price
- $0.2 / 1M tokens
- Output price
- $0.88 / 1M tokens
- Released
- 2026-02-05
Capabilities
- Streaming
- Playground
Provider pricing, discounts and data privacy
Compare effective provider prices, published discounts, regions, retention policies, training use, compliance, and privacy links by service tier.
Standard service tier
3 available providers · tier input average $0.343333 / 1M tokens · tier output average