Deepseek: Deepseek R1 Distill Qwen 14B
deepseek/deepseek-r1-distill-qwen-14b
DeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 69.7 MATH-500 pass@1: 93.9 CodeForces Rating: 1481 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
Model specifications
- Input
- text
- Output
- text
- Context
- 32,768 tokens
- Max output
- 16,400 tokens
- Input price
- $0 / 1M tokens
- Output price
- $0 / 1M tokens
- Released
- 2026-02-18
Capabilities
- Streaming
- Function calling
- JSON mode
- Playground
Frequently asked questions
- What is Deepseek: Deepseek R1 Distill Qwen 14B?
- DeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 69.7 MATH-500 pass@1: 93.9 CodeForces Rating: 1481 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
- How much does Deepseek: Deepseek R1 Distill Qwen 14B cost?
- Input costs start at $0 / 1M tokens and output costs start at $0 / 1M tokens. Provider-level prices vary by service tier.
- What is the context length of Deepseek: Deepseek R1 Distill Qwen 14B?
- Deepseek: Deepseek R1 Distill Qwen 14B supports a 32,768 token context window and up to 16,400 output tokens.
- What capabilities does Deepseek: Deepseek R1 Distill Qwen 14B support?
- Deepseek: Deepseek R1 Distill Qwen 14B supports Streaming, Function calling, JSON mode, Playground.