Multimodal Reference Video Standard Edition for High-Definition Finished Videos
doubao-seedance-2-0-260128 is ByteDance Seedance 2.0 Standard Edition, suitable for combining text, character images, and audio-video references into short videos. It can bring static images to life while using reference materials to constrain subjects, actions, and camera pacing, with support for sound generation. When you need high-definition output, character references, and short-shot creation, this is a version worth prioritizing.
Up to 9 images; up to 3 audio clips and 3 videos each
Sound control
Enable generate_audio for sound generation; disabled by default
Delivery method
Video link; supports asynchronous tasks, callbacks, and final-frame return
The above describes this platform's supported invocation scope for this model. Native capabilities and material requirements for each creation mode should be understood separately.
Core Capabilities
Learn what doubao-seedance-2-0-260128 can bring to your work.
Separate Control of Subject Reference and Scene Creation
Use reference_image to provide a person or subject, then specify the scene, clothing, actions, and camera with text, allowing the reference image to carry identity and appearance cues rather than fixing the entire composition. If you want the video to begin moving from the original image composition, use first_frame instead, which is suitable for turning portraits, product images, or illustrations into dynamic shots.
Express Motion and Rhythm with Audio and Video
In addition to images, you can add reference_video and reference_audio to provide cues for motion, camera movement, sound, and rhythm. Text should specify which parts of each asset need to be referenced to avoid conflicting requirements from different assets; when you need a finished video with sound, enable generate_audio so visuals and sound enter the same generation process.
From Preview Formats to High-Definition Delivery
The standard version offers resolution options from 480p to 4k and supports landscape, portrait, square, and ultra-wide formats. You can first validate the subject and motion, then choose a higher resolution to produce delivery shots; use return_last_frame to obtain the final frame, which also makes it easier to preview the ending or prepare visual assets for the next creation.
Use Cases
Start with specific tasks to find where the model can make an impact.
Short Product Showcase Shots
Input a product image and a shot description, such as a slow push-in, sideways movement, or highlighting material details, to generate standalone clips for advertising edits. Use first-frame mode when you need to preserve the starting layout, and reference mode when you need to present the same subject in a new scene, then choose landscape, portrait, or square framing based on placement.
Character Shorts Across Scenes
Provide clear reference images of people you own or are authorized to use, and describe the character's behavior in indoor, street, or natural environments to create multiple short shots of the same character. The reference image handles appearance cues, while text handles the scene and actions; this is suitable for character concept presentations and storyboard sequences. Before delivery, check each segment to ensure the character's appearance and actions meet requirements.
Motion and Sound Proof of Concept
Combine example motion videos, rhythmic audio, and scene prompts to explore visual expression for dance, sports, or scenes with ambient sound. First clarify whether you want to draw from the motion, camera work, or sound, generate short videos for the creative team to discuss, then download the videos for the editing workflow, avoiding treating every element in a reference asset as a target that must be copied.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Standard Edition or Fast, Mini
When 1080p or 4k output is needed, prioritize this Standard Edition; 2.0 Fast and Mini support 480p and 720p resolutions, making them more suitable for previews and lightweight creation. All three support character and audio/video references, but they are different model variants. Do not assume that simply replacing the name will result in exactly the same parameter combinations and output performance.
HD Short Shots or Longer Editing Tasks
This model is suitable for 4–15-second HD short shots and multimodal reference generation. If you need up to 30 seconds, audio-only references, more assets, or to edit and extend existing videos, choose Seedance 2.5. The key trade-off between the two is HD output and workflow type, rather than interpreting the newer version as having higher specifications in every respect.
Get Started
From a small-scale task to production integration.
01
Prepare Tasks and Materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Playground
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limits
Understand output quality and capability scope before formal use.
First-frame, first-and-last-frame, and all-modal reference modes are mutually exclusive. first_frame and last_frame cannot be used together with various reference tags; in multimodal creation, text can describe an image as the intended starting or ending frame, but when the first and last frames need to be locked, use the dedicated first-and-last-frame mode instead.
Reference audio must be wav or mp3, 2–15 seconds per clip and no more than 15 MB; reference video must be mp4 or mov, likewise 2–15 seconds per clip. The total duration of audio and video must each not exceed 15 seconds. Trim materials first to avoid failures during processing.
This model is intended for generation; do not treat editing, extension, web-connected tools, or mov output selection as its features. People and character materials must be owned or authorized, and photos should be clear, front-facing, and unobstructed; reference materials guide generation and do not mean that every action and detail will be reproduced exactly.
Frequently Asked Questions
Answers to common questions about using doubao-seedance-2-0-260128.
What is the difference between a reference image and a first-frame image?
reference_image provides clues about a person or subject, while the scene and action can still be redefined with text; first_frame makes the video start from the specified image. To place the same person in a new environment, choose a reference image; to make an existing photo composition come to life directly, choose the first frame. Do not mix the two tags.
How do I generate a video with sound?
Set generate_audio to true and describe the desired sound or ambient audio in the prompt to request a video with sound; this option is disabled by default. If you need reference audio, you must use the audio_url object and the reference_audio tag, while also providing reference_image or reference_video; you cannot upload only reference audio. Sound generation and audio reference are different controls, and uploading audio cannot replace the setting for generating sound; when not using reference audio, you can also describe the desired sound in text.
Can this version generate 30-second videos or extend existing videos?
This version generates videos lasting 4–15 seconds, and you can also use duration=-1 to determine the duration automatically. For up to 30 seconds or editing and extending existing videos, choose Seedance 2.5; adding reference_video is a reference generation method and does not mean editing or continuing the original video.
How should I specify image URLs when making a call?
Submit the exact model ID and content array to POST /seedance/videos. Image items use an image_url object, for example containing url inside it, rather than writing the address directly as a string; use type, role, and text prompts to describe the purpose, and set resolution and duration using top-level fields.
How do I retrieve the finished video after submission?
You can set async=true to obtain a task_id and then query the task, or provide callback_url to receive a completion notification. After obtaining the completed result, read data.video_url to download the video; do not treat a successful submission as meaning the finished video is ready. If you need a final-frame image, you can also enable return_last_frame.
Model information · Updated: 2026-10-01. See the API and pricing sections for call parameters and billing rules.