A short-video creation model focused on visual quality and consistency
kling-v2-1-master is Kuaishou Kling V2.1 Master video model, suitable for short-form creation that prioritizes visual quality and visual consistency. It can build shots from text descriptions or start animations from a first-frame image, generating 5-second or 10-second videos. On this platform, you can also use the talking photo feature to turn portrait photos and existing audio into lip-synced videos.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation methods
Text-to-video, first-frame image-to-video
Video duration
5 seconds or 10 seconds
Video aspect ratios
16:9, 9:16, 1:1
Prompt controls
prompt, negative_prompt; cfg_scale range 0–1
Generation results
Video link, video ID, task ID, duration, and status
Talking photo inputs
Portrait photo and audio link; supported audio formats: mp3/wav/m4a/aac, ≤5MB
The above are the available specifications for the corresponding entry point on this platform; talking photos are created through animation and lip synchronization and are not native model audio.
Core capabilities
Prioritize visual quality and consistency
V2.1 Master is positioned with an emphasis on visual quality and consistency, making it suitable for focusing on subject presentation, scene atmosphere, and the overall look of a shot. When creating, use text to clearly define the subject, action, environment, and lighting, then refine descriptions around the same shot, compare different results, and select clips suitable for editing and presentation.
Start from a text concept or existing image
When no ready-made visual material is available, use text-to-video to explore shots; when you already have product images, portraits, or scene images, use first-frame image-to-video to begin from a specified frame. The two methods serve concept exploration and asset animation respectively. Prompts should focus on describing what happens next, rather than merely repeating the image content.
Make portrait photos speak with existing audio
The talking photo feature accepts a portrait photo and recorded audio, first generating photo animation and then completing lip synchronization. Use prompts to add expression or movement requirements, producing the final talking-head video and an intermediate animated video. It is suitable for workflows with existing audio assets that need to quickly create short character-expression clips.
Applicable Scenarios
Turn Product Images into Showcase Clips
Use a product image as the first frame, describe the desired actions, environmental changes, and display highlights, and create short videos suitable for product introductions or social content. Choose a landscape, portrait, or square aspect ratio based on the placement, and deliver video assets ready for the editing workflow; this model does not use an end-frame image to specify the final composition.
Script Shot and Atmosphere Previsualization
Break a script into individual shots, specify the subject, action, and scene for each segment, and use text-to-video to create visual drafts. 5-second clips can be used for brief action demonstrations, while 10-second clips can accommodate more complete process expression, producing previsualization materials for discussing composition, pacing, and atmosphere rather than completing an entire film in one go.
Photo Talking Heads and Character Lines
Prepare a clear front-facing photo of one person and a short audio clip to create videos for introductions, greetings, or character dialogue. The audio is recommended to be no longer than the selected video duration, and prompts can add the desired expressions and actions. The final video is used for talking-head delivery, while intermediate animations can be used to check the photo's dynamic performance.
How to Choose This Model
For a Clear Quality Focus, Prioritize V2.1 Master
When choosing between V2 Master and V2.1 Master, consider V2.1 Master as the option more focused on quality and consistency, and compare actual results using the same prompt or first frame. Neither should be understood as supporting all advanced controls simply because of the Master name. If end-frame constraints are needed, consider V2.5 Turbo pro, which supports this capability.
Choose Another Option When You Need Audio or Multi-Asset Editing
For standard video creation requiring synchronized audio, consider V2.6 pro or V3; for native 4K, consider the corresponding modes in the V3 series. If the task involves multi-image references, reference videos, or editing existing videos, choose O1 or V3 Omni. For short talking-head videos using existing photos and audio, you can continue using this model's talking photo feature.
Getting Started
Define the starting composition for the person or product
For text prompts, choose text2video; for photo animation, choose image2video and provide start_image_url. Clearly describe the person's actions, the product's appearance, and the camera intent separately.
Choose a five- or ten-second clip
Explicitly specify model=kling-v2-1-master for /kling/videos, use this model's single mode, and start with 5 seconds; choose 10 seconds when a complete action is needed. Image-to-video provides only the first frame.
Check appearance and continuity segment by segment
Save the asynchronous task_id and retrieve the video through queries or callbacks; compare it with the original image to check the person, clothing, and background. Standard video generation does not include native audio, so add sound in post-production; photo lip-sync requires a separate workflow using a photo and existing audio.
Trial Suggestion: Character Motion Storyboard
Input and goal
Use a photo of a person standing by a window as the first frame. The person turns toward the camera and smiles, while the camera slowly moves closer, preserving the clothing and indoor layout.
Review and next steps
Check the face, clothing, and motion; this model uses the first frame and does not support workflows extended through multi-image lists, end frames, or native audio.
Usage Limitations
Standard video generation does not support end-frame images, structured camera_control camera movement controls, or native synchronized audio. You can describe camera intent in the prompt, but such descriptions should not be treated as precisely configurable camera movement parameters, and tasks should not be designed around a first-to-last-frame interpolation workflow.
This model organizes creation around 5- or 10-second clips. It does not support native 4K mode and does not provide Omni multi-image references or reference-video editing. For longer content, split it into shots before editing; do not use multiple reference assets or extended videos as the default creative approach.
Photo lip-sync requires existing audio; entering dialogue does not automatically generate sound. Photos should preferably be clear, feature a single person facing forward, and use audio files no larger than 5MB with a duration that matches the video as closely as possible; profile views, occlusions, or overly long audio are not ideal source material.
Frequently Asked Questions
How should I choose between V2.1 Master and V2 Master?
V2.1 Master is better suited as an option focused on visual quality and consistency. It is recommended to compare the actual results of both using the same prompt or first frame, rather than looking only at the version names. Both are used here for 5- or 10-second short clips, which does not imply additional control capabilities.
Can image-to-video specify both a first frame and a last frame at the same time?
This model uses first-frame image-to-video: submit start_image_url and describe the subsequent actions with a prompt; end_image_url last-frame constraints are not supported. If you must specify the ending image, you can choose a model that supports first and last frames, such as V2.5 Turbo pro, and prepare assets according to the corresponding creation workflow.
What is the difference between regular videos and talking photos?
Regular videos generate dynamic visuals from text or a first-frame image; this model does not natively generate synchronized audio. Talking photos, however, require a photo and existing audio, generating a spoken video through animation and lip synchronization; the audio is an input asset, not a voice automatically created by the model.
How do I submit a video task for this model?
Call POST /kling/videos, explicitly set model to kling-v2-1-master, choose text2video or image2video, and submit a prompt with a duration of 5 or 10 seconds. Image-to-video also requires a first-frame image link; after completion, retrieve the generated video through video_url.
How can I track generation results when producing in batches?
You can set async=true, first save the returned task_id, then query task progress; you can also configure callback_url to receive completion notifications. After obtaining results, associate video links and statuses by task ID. Photo narration can also use the asynchronous approach and returns the final video along with intermediate animation links.