ByteDance: Seedance 2.5 (Reference to Video) API Reference

About

Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.

1. Calling the API

Setup your API Key

To begin using Infron, you first need to create an account and generate your API Key.

Once you have your API key, set it as an environment variable in your runtime environment. This allows your applications to securely access Infron’s API.

export INFRON_KEY="YOUR_API_KEY"

Submit a request

For long-running requests, you can poll for results.

curl https://media.onerouter.pro/v1/videos/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "bytedance/seedance-2.5/reference-to-video",
  "prompt": "Style: Hybrid visual style — photorealistic, documentary-level environment combined with stylized 3D animated characters. The subject and the fan are fully 3D animated characters seamlessly composited into a live-action realistic world. Single continuous unbroken shot from a handheld camera within a dense crowd. Natural micro-shake, eye-level perspective.
Character Style: The subject from @Image1 is rendered as a polished 3D animated character with stylized proportions, soft subsurface skin shading, expressive features, and clean rim lighting — while maintaining a perfectly consistent face and the exact outfit from the reference image. The fan they interact with is also a 3D animated character in the same rendering style. Both characters retain cinematic CG quality with realistic interaction with the surrounding light (flash bounces, streetlight highlights, shadow casting on real ground).
Lighting & Environment: Fully photorealistic. Nighttime at an upscale event in New York City. Illuminated by real streetlights and camera flashes. Mixed reflections on polished surfaces (phones, cars), soft realistic shadows, and a slight atmospheric haze for depth. The crowd, barricades, hotel facade, SUVs, and street are all live-action realistic — only the two main characters are stylized 3D.
Subject: The 3D animated subject maintains a calm, controlled presence with a subtle, confident smile, perfectly matching the face and outfit from @Image1.
Action Sequence: The shot begins completely immersed in a restless, chaotic realistic crowd behind barricades. The view is partially obscured by real people raising smartphones to record. As the camera lifts slightly above shoulder level, the 3D animated subject exits a luxury hotel in the background. Bright media flashes erupt, illuminating the CG character against the realistic environment. Real security personnel step into frame, pushing the crowd back, causing the camera to shake naturally. Through shifting gaps in the crowd, the animated subject walks forward clearly into center frame. The subject pauses to interact with a 3D animated fan, leaning in briefly for a selfie while giving a calm, controlled wave. The camera pans to follow as a luxury convoy of three premium black SUVs (photorealistic) pulls up. A real security guard opens the back door of the middle SUV. The animated subject steps inside, rolls down the window to wave one last time, and the vehicles begin to pull away as the realistic crowd jumps to capture the moment.
Audio: Loud, chaotic crowd cheering and whistling. Overlapping voices shouting the subject's name. A barrage of rapid camera shutter clicks. Distant New York City sirens and traffic. The rustling of heavy fabric and footsteps. The deep, heavy bass of an SUV engine idling and pulling away.",
  "image_urls": ["https://storage.googleapis.com/infron_gcs/quick-upload/2026/08/10/20260810-160755-d1d01d5c-1.png"],
  "video_urls": [],
  "audio_urls": [],
  "aspect_ratio": "adaptive",
  "duration": "-1",
  "resolution": "720p",
  "generate_audio": true,
  "n": 1,
  "seed": 123456
}
EOF

Response with queue task id

The response will look like this:

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "20260810080823640191553lDkmm6gD",
    "object": "video",
    "model": "bytedance/seedance-2.5/reference-to-video",
    "status": "created",
    "urls": {
      "query": "https://media.onerouter.pro/v1/videos/tasks/20260810080823640191553lDkmm6gD"
    },
    "created_at": "2026-08-10 08:08:23.70"
  }
}

Fetch request status

You can fetch the status of a request to check if it is completed or still in progress.

curl https://media.onerouter.pro/v1/videos/tasks/{task_id} \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" 

# For example
curl https://media.onerouter.pro/v1/videos/tasks/20260810080823640191553lDkmm6gD \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json"

Response with queue task status query

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "20260810080823640191553lDkmm6gD",
    "object": "video",
    "model": "bytedance/seedance-2.5/reference-to-video",
    "status": "in_progress",
    "fail_reason": "",
    "submit_time": 1786349304,
    "start_time": 1786349305,
    "finish_time": 0,
    "outputs": [],
    "created_at": "2026-08-10 08:08:24"
  }
}

Possible statuses:

  • created
  • in_progress
  • processing
  • completed
  • failed

created -> in_progress → processing → completed | failed

Poll every 1-2 seconds until status is "completed" or "failed"

Get the result (when status is "completed")

Once the request is completed, you can fetch the result. See the Output Schema for the expected result format.

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "20260810080823640191553lDkmm6gD",
    "object": "video",
    "model": "bytedance/seedance-2.5/reference-to-video",
    "status": "completed",
    "fail_reason": "",
    "submit_time": 1786349304,
    "start_time": 1786349305,
    "finish_time": 1786349602,
    "outputs": [
      "https://infron.ai/cdn/media/video/video_v2/output/output_e01a8536-a52d-4377-a942-8df7e5461541.mp4"
    ],
    "usage": {
      "completion_tokens": 324900,
      "duration": 15,
      "generate_audio": true,
      "output_count": 1,
      "prompt_tokens": 0,
      "request_num": 1,
      "resolution": "720p",
      "tokens": 324900,
      "total_tokens": 324900,
      "video_duration": 15
    },
    "cost": {
      "total_cost": 3.47643,
      "cost_details": {
        "prompt_cost": 0,
        "completion_cost": 0,
        "image_cost": 0,
        "video_cost": 3.47643,
        "audio_cost": 0,
        "native_web_search_cost": 0,
        "plugin_web_search_cost": 0,
        "tools_cost": 0,
        "prompt_cache_read_cost": 0,
        "prompt_cache_write_cost": 0,
        "prompt_cache_write_5_min": 0,
        "prompt_cache_write_1_h": 0,
        "reasoning_cost": 0,
        "discount_rate": 1,
        "is_byok": false,
        "byok_cost": 0
      }
    },
    "created_at": "2026-08-10 08:08:24"
  }
}

2. Schema

Input

ParameterTypeRequiredDefaultRangeDescription
modelstringYes--Explore all modelsModel name.
promptstringYes----Text prompt for generation; Positive text prompt.
image_urlsarrayNo--[0,30]Image URLs to use as a reference for the generation.
video_urlsarrayNo--[0,10]Video URLs to use as a reference for the generation.
audio_urlsarrayNo--[0,10]Audio URLs to use as a reference for the generation.
aspect_ratiostringNoadaptive21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptiveAspect ratio of the generated video
durationstringNo-1-1, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30Duration of the generated video
resolutionstringNo720p480p, 720p, 1080pResolution of the generated video
generate_audiobooleanNotruetrue, falseWhether to generate audio.
nintegerNo1==1Number of videos to generate.
seedintegerNo-->=1The random seed to use for the generation.

Example request

curl https://media.onerouter.pro/v1/videos/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "bytedance/seedance-2.5/reference-to-video",
  "prompt": "Style: Hybrid visual style — photorealistic, documentary-level environment combined with stylized 3D animated characters. The subject and the fan are fully 3D animated characters seamlessly composited into a live-action realistic world. Single continuous unbroken shot from a handheld camera within a dense crowd. Natural micro-shake, eye-level perspective.
Character Style: The subject from @Image1 is rendered as a polished 3D animated character with stylized proportions, soft subsurface skin shading, expressive features, and clean rim lighting — while maintaining a perfectly consistent face and the exact outfit from the reference image. The fan they interact with is also a 3D animated character in the same rendering style. Both characters retain cinematic CG quality with realistic interaction with the surrounding light (flash bounces, streetlight highlights, shadow casting on real ground).
Lighting & Environment: Fully photorealistic. Nighttime at an upscale event in New York City. Illuminated by real streetlights and camera flashes. Mixed reflections on polished surfaces (phones, cars), soft realistic shadows, and a slight atmospheric haze for depth. The crowd, barricades, hotel facade, SUVs, and street are all live-action realistic — only the two main characters are stylized 3D.
Subject: The 3D animated subject maintains a calm, controlled presence with a subtle, confident smile, perfectly matching the face and outfit from @Image1.
Action Sequence: The shot begins completely immersed in a restless, chaotic realistic crowd behind barricades. The view is partially obscured by real people raising smartphones to record. As the camera lifts slightly above shoulder level, the 3D animated subject exits a luxury hotel in the background. Bright media flashes erupt, illuminating the CG character against the realistic environment. Real security personnel step into frame, pushing the crowd back, causing the camera to shake naturally. Through shifting gaps in the crowd, the animated subject walks forward clearly into center frame. The subject pauses to interact with a 3D animated fan, leaning in briefly for a selfie while giving a calm, controlled wave. The camera pans to follow as a luxury convoy of three premium black SUVs (photorealistic) pulls up. A real security guard opens the back door of the middle SUV. The animated subject steps inside, rolls down the window to wave one last time, and the vehicles begin to pull away as the realistic crowd jumps to capture the moment.
Audio: Loud, chaotic crowd cheering and whistling. Overlapping voices shouting the subject's name. A barrage of rapid camera shutter clicks. Distant New York City sirens and traffic. The rustling of heavy fabric and footsteps. The deep, heavy bass of an SUV engine idling and pulling away.",
  "image_urls": ["https://storage.googleapis.com/infron_gcs/quick-upload/2026/08/10/20260810-160755-d1d01d5c-1.png"],
  "video_urls": [],
  "audio_urls": [],
  "aspect_ratio": "adaptive",
  "duration": "-1",
  "resolution": "720p",
  "generate_audio": true,
  "n": 1,
  "seed": 123456
}
EOF

Output

ParameterTypeRangeDescription
codeinteger--Status code of the request.
messagestring--Status message of the request.
dataobject--Data of the video task.
data.task_idstring--ID of the video task.
data.objectstring--Object type of the video task.
data.modelstring--Model name of the video task.
data.statusstringcreated, in_progress, processing, completed, failedStatus of the video task.
data.fail_reasonstring--Fail reason of the video task.
data.submit_timeinteger--Timestamp of the video task submission.
data.start_timeinteger--Timestamp of the video task start.
data.finish_timeinteger--Timestamp of the video task finish.
data.outputsarray--Outputs of the video task.
data.created_atstring--Timestamp of the video task creation.

Example output

{
  "code": 200,
  "message": "success",
  "data": {
    "task_id": "20260810080823640191553lDkmm6gD",
    "object": "video",
    "model": "bytedance/seedance-2.5/reference-to-video",
    "status": "completed",
    "fail_reason": "",
    "submit_time": 1786349304,
    "start_time": 1786349305,
    "finish_time": 1786349602,
    "outputs": [
      "https://infron.ai/cdn/media/video/video_v2/output/output_e01a8536-a52d-4377-a942-8df7e5461541.mp4"
    ],
    "usage": {
      "completion_tokens": 324900,
      "duration": 15,
      "generate_audio": true,
      "output_count": 1,
      "prompt_tokens": 0,
      "request_num": 1,
      "resolution": "720p",
      "tokens": 324900,
      "total_tokens": 324900,
      "video_duration": 15
    },
    "cost": {
      "total_cost": 3.47643,
      "cost_details": {
        "prompt_cost": 0,
        "completion_cost": 0,
        "image_cost": 0,
        "video_cost": 3.47643,
        "audio_cost": 0,
        "native_web_search_cost": 0,
        "plugin_web_search_cost": 0,
        "tools_cost": 0,
        "prompt_cache_read_cost": 0,
        "prompt_cache_write_cost": 0,
        "prompt_cache_write_5_min": 0,
        "prompt_cache_write_1_h": 0,
        "reasoning_cost": 0,
        "discount_rate": 1,
        "is_byok": false,
        "byok_cost": 0
      }
    },
    "created_at": "2026-08-10 08:08:24"
  }
}

3. Detailed Pricing

ResolutionApprox. Price
720p with audio~$0.231 / second
480p with audio~$0.103 / second
GenerationTokensCost
5s at 720p 16:9~108,000~$1.16
10s at 720p 16:9~216,000~$2.31
30s at 720p 16:9~648,000~$6.93
5s at 480p 16:9~48,000~$0.51
10s at 480p 16:9~96,000~$1.03
30s at 480p 16:9~288,000~$3.08