All models

kling-v2-1-master

KuaishouVideo
Get your API key
kling-v2-1-master

Short Video Creation Model for Visual Quality and Consistency

kling-v2-1-master is Kuaishou Kling V2.1 Master video model, suitable for short-form creation that prioritizes visual quality and consistency. It can construct shots from text descriptions or use a first-frame image to start animation, generating 5-second or 10-second videos. On this platform, you can also use the talking photo feature to create lip-synced videos from portrait photos and existing audio.

KuaishouModel Brand
VideoModel Type
VideoTask Capability
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelkling-v2-1-master

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Clarify capacity, inputs and outputs, and invocation methods before choosing a model.

Creation Methods
Text-to-video, first-frame image-to-video
Video Duration
5 seconds or 10 seconds
Video Aspect Ratios
16:9, 9:16, 1:1
Prompt Control
prompt, negative_prompt; cfg_scale range 0–1
Generation Results
Video link, video ID, task ID, duration, and status
Talking Photo Input
Portrait photo and audio link; supported audio formats: mp3/wav/m4a/aac, ≤5MB

The above are the available specifications for the corresponding entry point on this platform; talking photos are created through animation and lip synchronization and do not include the model's native audio.

Core Capabilities

Learn what kling-v2-1-master can bring to your work.

Prioritize visual quality and consistency

V2.1 Master is positioned with an emphasis on visual quality and consistency, making it suitable for focusing on subject performance, scene atmosphere, and the overall look of the shot. When creating, use text to clearly specify the subject, action, environment, and lighting, then refine the description around the same shot, compare different results, and select clips suitable for editing and presentation.

Start from a text concept or an existing image

When no visual material is available, use text-to-video to explore shots; when you already have product images, portraits, or scene images, use first-frame image-to-video to begin from a specified image. The two methods serve concept exploration and animating existing assets respectively, and prompts should focus on describing what happens next rather than simply repeating the image content.

Make portrait photos speak with existing audio

The talking photo feature accepts a portrait photo and recorded audio, first generating photo animation and then completing lip-syncing. Use prompts to add facial expression or motion requirements to obtain the final talking-head video and an intermediate animated video, making it suitable for workflows with existing audio assets that need to quickly produce short character expressions.

Use Cases

Start from specific tasks and find where the model can be effective.

Turn product images into showcase clips

Use a product image as the first frame, describe the desired action, environmental changes, and presentation focus, and create short videos suitable for product introductions or social content. Choose landscape, portrait, or square aspect ratios based on the placement, and deliver video assets ready for the editing workflow; this model does not use an end-frame image to specify the final composition.

Preview script shots and atmosphere

Break a script into individual shots, specify the subject, action, and scene for each segment, and use text-to-video to create visual drafts. Five-second clips can be used for brief action demonstrations, while ten-second clips can accommodate more complete process expression, producing preview assets for discussing composition, pacing, and atmosphere rather than completing an entire film in one pass.

Photo talking-head videos and short character lines

Prepare a clear front-facing portrait photo of one person and a short audio clip to create videos for introductions, greetings, or character dialogue. The audio is recommended to be no longer than the selected video duration, and prompts can add desired expressions and actions. The final video is used for talking-head delivery, while the intermediate animation can be used to check the photo's dynamic performance.

How to choose this model

Choose based on task complexity, input materials, and expected results.

For a clear quality focus, prioritize V2.1 Master

When choosing between V2 Master and V2.1 Master, consider V2.1 Master as a candidate with greater emphasis on quality and consistency, and compare actual results using the same prompt or first frame. Neither should be understood as supporting all advanced controls simply because of the Master name. If you want to use end-frame constraints, consider V2.5 Turbo pro, which supports this capability.

Choose another option when sound or multi-asset editing is needed

When standard video creation requires synchronized audio, consider V2.6 pro or V3; when native 4K is needed, consider the corresponding mode in the V3 series. For tasks involving multi-image references, reference videos, or editing existing videos, choose O1 or V3 Omni. For short talking-head videos using existing photos and audio, you can continue using this model's talking photo feature.

Get started

From a small-scale task to formal integration.

01

Prepare the task and materials

Define the goal, required inputs, and output requirements, using real business examples as a starting point.

02

Try it in the API debugging area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate according to the API documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage limitations

Before formal use, understand the output quality and capability scope.

  • Standard video generation does not support end-frame images, structured camera_control camera movement controls, or native synchronized audio. You can describe camera intentions in the prompt, but such descriptions cannot be treated as precisely configurable camera movement parameters, nor should tasks be designed around a first-to-last-frame interpolation workflow.
  • This model organizes creation into 5-second or 10-second clips, does not support a native 4K mode, and does not provide Omni multi-image references or reference video editing. Longer content should preferably be split into shots before editing; do not use multiple reference assets or extended video as the default creation approach.
  • Photo narration requires existing audio; entering a script does not automatically generate sound. Photos should preferably be clear, feature one person, and show a front-facing face; audio files must not exceed 5MB, and their duration should match the video where possible. Profile views, obstructions, or overly long audio are not recommended as first-choice materials.

Frequently Asked Questions

Answers to common questions about using kling-v2-1-master.

How should I choose between V2.1 Master and V2 Master?

V2.1 Master is better suited as a candidate for image quality and consistency. It is recommended to compare the actual results of both using the same prompt or first frame, rather than judging only by the version name. Both are used here for 5- or 10-second clips, which does not imply additional control capabilities.

Can image-to-video specify both a first frame and a last frame?

This model uses first-frame image-to-video: submit start_image_url and describe subsequent actions with a prompt; end_image_url last-frame constraints are not supported. If you must specify the ending image, choose a model that supports first and last frames, such as V2.5 Turbo pro, and prepare materials according to the corresponding creation workflow.

What is the difference between regular video and talking photos?

Regular video generates dynamic visuals from text or a first-frame image; this model does not natively generate synchronized audio. Talking photos, on the other hand, require a photo and existing audio, generating a spoken video through animation and lip synchronization; the audio is input material, not a voice automatically created by the model.

How do I submit a video task for this model?

Call POST /kling/videos, explicitly set model to kling-v2-1-master, choose text2video or image2video, and submit a prompt with a duration of 5 or 10 seconds. Image-to-video also requires a first-frame image link; after completion, retrieve the generated video through video_url.

How can I track generation results for batch production?

You can set async=true, save the returned task_id first, then query task progress; you can also configure callback_url to receive completion notifications. After obtaining results, associate video links and statuses by task ID. Talking photos can likewise use an asynchronous approach and return the final video and intermediate animation links.

Model information · Updated: 2026-10-01. For API parameters and billing rules, see the API and pricing sections.