Deepseek: Deepseek R1 Distill Qwen 14B

deepseek/deepseek-r1-distill-qwen-14b

DeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 69.7 MATH-500 pass@1: 93.9 CodeForces Rating: 1481 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

Model specifications

Input
text
Output
text
Context
32,768 tokens
Max output
16,400 tokens
Input price
$0 / 1M tokens
Output price
$0 / 1M tokens
Released
2026-02-18

Capabilities

  • Streaming
  • Function calling
  • JSON mode
  • Playground

Frequently asked questions

What is Deepseek: Deepseek R1 Distill Qwen 14B?
DeepSeek R1 Distill Qwen 14B is a distilled large language model based on Qwen 2.5 14B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models. Other benchmark results include: AIME 2024 pass@1: 69.7 MATH-500 pass@1: 93.9 CodeForces Rating: 1481 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.
How much does Deepseek: Deepseek R1 Distill Qwen 14B cost?
Input costs start at $0 / 1M tokens and output costs start at $0 / 1M tokens. Provider-level prices vary by service tier.
What is the context length of Deepseek: Deepseek R1 Distill Qwen 14B?
Deepseek: Deepseek R1 Distill Qwen 14B supports a 32,768 token context window and up to 16,400 output tokens.
What capabilities does Deepseek: Deepseek R1 Distill Qwen 14B support?
Deepseek: Deepseek R1 Distill Qwen 14B supports Streaming, Function calling, JSON mode, Playground.

Browse all AI models