Deepseek: Deepseek R1 Distill Qwen 32B
deepseek/deepseek-r1-distill-qwen-32b
DeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.\n\nOther benchmark results include:\n\n- AIME 2024 pass@1: 72.6\n- MATH-500 pass@1: 94.3\n- CodeForces Rating: 1691\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
Model specifications
- Input
- text
- Output
- text
- Context
- 131,072 tokens
- Max output
- 32,800 tokens
- Input price
- $0 / 1M tokens
- Output price
- $0 / 1M tokens
- Released
- 2026-02-18
Capabilities
- Streaming
- Function calling
- JSON mode
- Playground
Frequently asked questions
- What is Deepseek: Deepseek R1 Distill Qwen 32B?
- DeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.\n\nOther benchmark results include:\n\n- AIME 2024 pass@1: 72.6\n- MATH-500 pass@1: 94.3\n- CodeForces Rating: 1691\n\nThe model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
- How much does Deepseek: Deepseek R1 Distill Qwen 32B cost?
- Input costs start at $0 / 1M tokens and output costs start at $0 / 1M tokens. Provider-level prices vary by service tier.
- What is the context length of Deepseek: Deepseek R1 Distill Qwen 32B?
- Deepseek: Deepseek R1 Distill Qwen 32B supports a 131,072 token context window and up to 32,800 output tokens.
- What capabilities does Deepseek: Deepseek R1 Distill Qwen 32B support?
- Deepseek: Deepseek R1 Distill Qwen 32B supports Streaming, Function calling, JSON mode, Playground.