ByteDance: Seedance 2.0 Mini (Reference to Video) API Reference

About

Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.

1. Calling the API

Setup your API Key

To begin using Infron, you first need to create an account and generate your API Key.

Once you have your API key, set it as an environment variable in your runtime environment. This allows your applications to securely access Infron’s API.

export INFRON_KEY="YOUR_API_KEY"

Submit a request

For long-running requests, you can poll for results.

curl https://media.onerouter.pro/v1/videos/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "bytedance/seedance-2.0/mini/reference-to-video",
  "prompt": "Style: Hybrid visual style — photorealistic, documentary-level environment combined with stylized 3D animated characters. The subject and the fan are fully 3D animated characters seamlessly composited into a live-action realistic world. Single continuous unbroken shot from a handheld camera within a dense crowd. Natural micro-shake, eye-level perspective.
Character Style: The subject from @Image1 is rendered as a polished 3D animated character with stylized proportions, soft subsurface skin shading, expressive features, and clean rim lighting — while maintaining a perfectly consistent face and the exact outfit from the reference image. The fan they interact with is also a 3D animated character in the same rendering style. Both characters retain cinematic CG quality with realistic interaction with the surrounding light (flash bounces, streetlight highlights, shadow casting on real ground).
Lighting & Environment: Fully photorealistic. Nighttime at an upscale event in New York City. Illuminated by real streetlights and camera flashes. Mixed reflections on polished surfaces (phones, cars), soft realistic shadows, and a slight atmospheric haze for depth. The crowd, barricades, hotel facade, SUVs, and street are all live-action realistic — only the two main characters are stylized 3D.
Subject: The 3D animated subject maintains a calm, controlled presence with a subtle, confident smile, perfectly matching the face and outfit from @Image1.
Action Sequence: The shot begins completely immersed in a restless, chaotic realistic crowd behind barricades. The view is partially obscured by real people raising smartphones to record. As the camera lifts slightly above shoulder level, the 3D animated subject exits a luxury hotel in the background. Bright media flashes erupt, illuminating the CG character against the realistic environment. Real security personnel step into frame, pushing the crowd back, causing the camera to shake naturally. Through shifting gaps in the crowd, the animated subject walks forward clearly into center frame. The subject pauses to interact with a 3D animated fan, leaning in briefly for a selfie while giving a calm, controlled wave. The camera pans to follow as a luxury convoy of three premium black SUVs (photorealistic) pulls up. A real security guard opens the back door of the middle SUV. The animated subject steps inside, rolls down the window to wave one last time, and the vehicles begin to pull away as the realistic crowd jumps to capture the moment.
Audio: Loud, chaotic crowd cheering and whistling. Overlapping voices shouting the subject's name. A barrage of rapid camera shutter clicks. Distant New York City sirens and traffic. The rustling of heavy fabric and footsteps. The deep, heavy bass of an SUV engine idling and pulling away.",
  "image_urls": ["https://storage.googleapis.com/infron_gcs/image%2Fplayground-inputs%2F15%2F20260602%2F9ff4954d-be08-4a67-b868-ba1857616c26.png"],
  "video_urls": [],
  "audio_urls": [],
  "aspect_ratio": "16:9",
  "duration": "15",
  "resolution": "720p",
  "generate_audio": true,
  "n": 1,
  "seed": 123456
}
EOF

Response with queue task id

The response will look like this:

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "40feb57984fd47dfab34caeecdc22be9",
    "object": "video",
    "model": "bytedance/seedance-2.0/mini/reference-to-video",
    "status": "created",
    "urls": {
      "query": "https://media.onerouter.pro/v1/videos/tasks/40feb57984fd47dfab34caeecdc22be9"
    },
    "created_at": "2026-06-27 10:06:30.29"
  }
}

Fetch request status

You can fetch the status of a request to check if it is completed or still in progress.

curl https://media.onerouter.pro/v1/videos/tasks/{task_id} \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" 

# For example
curl https://media.onerouter.pro/v1/videos/tasks/40feb57984fd47dfab34caeecdc22be9 \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json"

Response with queue task status query

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "40feb57984fd47dfab34caeecdc22be9",
    "object": "video",
    "model": "bytedance/seedance-2.0/mini/reference-to-video",
    "status": "in_progress",
    "fail_reason": "",
    "submit_time": 1782554790,
    "start_time": 1782554793,
    "finish_time": 0,
    "outputs": [],
    "created_at": "2026-06-27 10:06:30"
  }
}

Possible statuses:

  • created
  • in_progress
  • processing
  • completed
  • failed

created -> in_progress → processing → completed | failed

Poll every 1-2 seconds until status is "completed" or "failed"

Get the result (when status is "completed")

Once the request is completed, you can fetch the result. See the Output Schema for the expected result format.

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "40feb57984fd47dfab34caeecdc22be9",
    "object": "video",
    "model": "bytedance/seedance-2.0/mini/reference-to-video",
    "status": "completed",
    "fail_reason": "",
    "submit_time": 1782554790,
    "start_time": 1782554793,
    "finish_time": 1782554988,
    "outputs": [
      "https://storage.googleapis.com/infron_gcs/video%2Fvideo_v2%2Foutput%2Foutput_0504e8e2-cb03-465d-b0c8-c7372f120bdd.mp4"
    ],
    "usage": {
      "completion_tokens": 324900,
      "duration": 15,
      "generate_audio": true,
      "output_count": 1,
      "prompt_tokens": 0,
      "request_num": 1,
      "resolution": "720p",
      "tokens": 324900,
      "total_tokens": 324900,
      "video_duration": 15
    },
    "cost": {
      "total_cost": 1.13715,
      "cost_details": {
        "prompt_cost": 0,
        "completion_cost": 0,
        "image_cost": 0,
        "video_cost": 1.13715,
        "audio_cost": 0,
        "native_web_search_cost": 0,
        "plugin_web_search_cost": 0,
        "tools_cost": 0,
        "prompt_cache_read_cost": 0,
        "prompt_cache_write_cost": 0,
        "prompt_cache_write_5_min": 0,
        "prompt_cache_write_1_h": 0,
        "reasoning_cost": 0,
        "discount_rate": 1,
        "is_byok": false,
        "byok_cost": 0
      }
    },
    "created_at": "2026-06-27 10:06:30"
  }
}

2. Schema

Input

ParameterTypeRequiredDefaultRangeDescription
modelstringYes--Explore all modelsModel name.
promptstringYes----Text prompt for generation; Positive text prompt.
image_urlsarrayNo--[1,9]Image URLs to use as a reference for the generation.
video_urlsarrayNo--[1,3]Video URLs to use as a reference for the generation.
audio_urlsarrayNo--[1,3]Audio URLs to use as a reference for the generation.
aspect_ratiostringNo16:921:9, 16:9, 4:3, 1:1, 3:4, 9:16Aspect ratio of the generated video
durationstringNo54, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15Duration of the generated video
resolutionstringNo720p480p, 720pResolution of the generated video
generate_audiobooleanNotruetrue, falseWhether to generate audio.
nintegerNo1==1Number of videos to generate.
seedintegerNo-->=1The random seed to use for the generation.

Example request

curl https://media.onerouter.pro/v1/videos/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "bytedance/seedance-2.0/mini/reference-to-video",
  "prompt": "Style: Hybrid visual style — photorealistic, documentary-level environment combined with stylized 3D animated characters. The subject and the fan are fully 3D animated characters seamlessly composited into a live-action realistic world. Single continuous unbroken shot from a handheld camera within a dense crowd. Natural micro-shake, eye-level perspective.
Character Style: The subject from @Image1 is rendered as a polished 3D animated character with stylized proportions, soft subsurface skin shading, expressive features, and clean rim lighting — while maintaining a perfectly consistent face and the exact outfit from the reference image. The fan they interact with is also a 3D animated character in the same rendering style. Both characters retain cinematic CG quality with realistic interaction with the surrounding light (flash bounces, streetlight highlights, shadow casting on real ground).
Lighting & Environment: Fully photorealistic. Nighttime at an upscale event in New York City. Illuminated by real streetlights and camera flashes. Mixed reflections on polished surfaces (phones, cars), soft realistic shadows, and a slight atmospheric haze for depth. The crowd, barricades, hotel facade, SUVs, and street are all live-action realistic — only the two main characters are stylized 3D.
Subject: The 3D animated subject maintains a calm, controlled presence with a subtle, confident smile, perfectly matching the face and outfit from @Image1.
Action Sequence: The shot begins completely immersed in a restless, chaotic realistic crowd behind barricades. The view is partially obscured by real people raising smartphones to record. As the camera lifts slightly above shoulder level, the 3D animated subject exits a luxury hotel in the background. Bright media flashes erupt, illuminating the CG character against the realistic environment. Real security personnel step into frame, pushing the crowd back, causing the camera to shake naturally. Through shifting gaps in the crowd, the animated subject walks forward clearly into center frame. The subject pauses to interact with a 3D animated fan, leaning in briefly for a selfie while giving a calm, controlled wave. The camera pans to follow as a luxury convoy of three premium black SUVs (photorealistic) pulls up. A real security guard opens the back door of the middle SUV. The animated subject steps inside, rolls down the window to wave one last time, and the vehicles begin to pull away as the realistic crowd jumps to capture the moment.
Audio: Loud, chaotic crowd cheering and whistling. Overlapping voices shouting the subject's name. A barrage of rapid camera shutter clicks. Distant New York City sirens and traffic. The rustling of heavy fabric and footsteps. The deep, heavy bass of an SUV engine idling and pulling away.",
  "image_urls": ["https://storage.googleapis.com/infron_gcs/image%2Fplayground-inputs%2F15%2F20260602%2F9ff4954d-be08-4a67-b868-ba1857616c26.png"],
  "video_urls": [],
  "audio_urls": [],
  "aspect_ratio": "16:9",
  "duration": "15",
  "resolution": "720p",
  "generate_audio": true,
  "n": 1,
  "seed": 123456
}
EOF

Output

ParameterTypeRangeDescription
codeinteger--Status code of the request.
messagestring--Status message of the request.
dataobject--Data of the video task.
data.task_idstring--ID of the video task.
data.objectstring--Object type of the video task.
data.modelstring--Model name of the video task.
data.statusstringcreated, in_progress, processing, completed, failedStatus of the video task.
data.fail_reasonstring--Fail reason of the video task.
data.submit_timeinteger--Timestamp of the video task submission.
data.start_timeinteger--Timestamp of the video task start.
data.finish_timeinteger--Timestamp of the video task finish.
data.outputsarray--Outputs of the video task.
data.created_atstring--Timestamp of the video task creation.

Example output

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "40feb57984fd47dfab34caeecdc22be9",
    "object": "video",
    "model": "bytedance/seedance-2.0/mini/reference-to-video",
    "status": "completed",
    "fail_reason": "",
    "submit_time": 1782554790,
    "start_time": 1782554793,
    "finish_time": 1782554988,
    "outputs": [
      "https://storage.googleapis.com/infron_gcs/video%2Fvideo_v2%2Foutput%2Foutput_0504e8e2-cb03-465d-b0c8-c7372f120bdd.mp4"
    ],
    "usage": {
      "completion_tokens": 324900,
      "duration": 15,
      "generate_audio": true,
      "output_count": 1,
      "prompt_tokens": 0,
      "request_num": 1,
      "resolution": "720p",
      "tokens": 324900,
      "total_tokens": 324900,
      "video_duration": 15
    },
    "cost": {
      "total_cost": 1.13715,
      "cost_details": {
        "prompt_cost": 0,
        "completion_cost": 0,
        "image_cost": 0,
        "video_cost": 1.13715,
        "audio_cost": 0,
        "native_web_search_cost": 0,
        "plugin_web_search_cost": 0,
        "tools_cost": 0,
        "prompt_cache_read_cost": 0,
        "prompt_cache_write_cost": 0,
        "prompt_cache_write_5_min": 0,
        "prompt_cache_write_1_h": 0,
        "reasoning_cost": 0,
        "discount_rate": 1,
        "is_byok": false,
        "byok_cost": 0
      }
    },
    "created_at": "2026-06-27 10:06:30"
  }
}

3. Detailed Pricing

Pricing TypeUnitPrice
Video generation with video input, 480p / 720p1K tokens$0.0021
Video generation without video input, 480p / 720p1K tokensTBD
Video generation, 1080p-Not supported
Video generation, 4K-Not supported
Input TypeResolutionPrice
With video input480p / 720p$2.10 / 1M tokens
Without video input480p / 720pTBD
With / without video input1080pNot supported
With / without video input4KNot supported
ResolutionDurationEstimated Price
480p5 s$0.19–$0.42
480p10 s~$0.38–$0.84
480p15 s~$0.57–$1.26
720p5 s$0.41–$0.91
720p10 s~$0.82–$1.82
720p15 s~$1.23–$2.73
1080p5 sNot supported
1080p10 sNot supported
1080p15 sNot supported
4K5 sNot supported
4K10 sNot supported
4K15 sNot supported