A short-film generation model combining first/last-frame control and synchronized audio
kling-v3 is the standard video generation model in Kuaishou Kling 3.0, suitable for creating short films from text concepts or still images. It supports flexible whole-second durations, first/last-frame guidance, synchronized audio, and multiple aspect ratios, making it suitable for both advertising shots and social content, as well as previewing motion in scenes. When you need to directly edit existing videos or combine multiple reference materials, choose a different Omni model.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Creation methods
Text-to-video, image-to-video; outputs video links
Video duration
3–15 seconds, whole seconds
Generation modes
std / pro / 4k; 4k and camera control are mutually exclusive
Aspect ratios
16:9、9:16、1:1
Image guidance
First-frame image link; can be paired with a last-frame image link
Audio
Can be generated synchronously; generate_audio defaults to false
Creation controls
std/pro supports camera_control; cfg_scale is 0–1, with optional negative prompts
The above are the invocation specifications for kling-v3 on this platform. Native creation capabilities and available parameters for each mode should be understood separately.
Core Capabilities
Learn what kling-v3 can bring to your work.
Turn text concepts into dynamic shots
Text-to-video is suitable for designing scenes from scratch: use prompts to specify the subject, action, environment, lighting, and camera intent, then choose the duration and aspect ratio. Relevance strength and negative prompts can help express creative constraints, making this suitable for iteratively refining the same shot rather than cramming an entire complex storyline into one short clip.
Guide changes with opening and closing frames
Image-to-video uses a first-frame image to establish the visual starting point, and an end frame can be added to specify the final image. For product showcases, character actions, or scene transitions, this approach can turn static designs into dynamic assets. Prompts should focus on how motion unfolds between the two images and avoid contradicting the subjects and composition in the images.
Balance audio and camera movement by task
When a sound-enabled short clip is needed, synchronized audio generation can be enabled; when more explicit camera movement is needed, camera control can be used in std or pro mode. 4k mode is suitable for tasks that prioritize output clarity, but camera_control cannot be added at the same time, so image-quality settings and camera control should be weighed before creation.
Use Cases
Start with specific tasks to find where the model can be effective.
Product showcases and advertising assets
Input a static product image, describe the display action, background atmosphere, and camera changes, and generate product shots for subsequent editing. If a clear final composition already exists, an end frame can also be provided. It is recommended to clearly specify the featured subject and core action, complete individual shots first, and then assemble them into a complete advertisement.
Vertical content and event clips
Use the event theme, setting, and subject action as text input, select a 9:16 aspect ratio and an appropriate whole-second duration, and create vertical short-video assets. Enable generate_audio when sound is needed; when brand posters or visual designs already exist, use image-to-video to bring the design into dynamic expression.
Storyboard previews and creative comparisons
Turn storyboard descriptions or keyframes into short clips for discussing action pacing, scene atmosphere, and camera direction. Keep subject and scene descriptions consistent, adjust only one key factor per round, and make it easier to compare different options. The deliverable is video footage suitable for further editing, rather than completing all post-production in one pass.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Choose the standard version for general generation, and Omni for asset editing
If the task is “generate a shot from a piece of text” or “make an image move,” kling-v3 offers a more direct workflow, along with first-and-last-frame and camera movement options. If you need to use multiple assets as independent references, or add or remove elements from existing videos and modify their style, choose kling-v3-omni or kling-o1; the standard version's first-and-last-frame guidance is not equivalent to video editing.
Better suited than older versions for flexible durations and sound-enabled creation
Compared with most older models that only accept 5 or 10 seconds, kling-v3 lets you set shots from 3–15 seconds in whole-second increments, making it easier to match the pacing of short videos. It supports synchronized audio, while kling-v2-6 limits audio to pro. Existing older-version workflows can be retained; prioritize this model when you need flexible duration, 4K, or a combination of camera movement and audio.
Get Started
From a small-scale task to production integration.
01
Prepare the task and materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try it in the API debugging area
Open the trial page, confirm the parameters supported by this endpoint, then submit a small-scale task to review the results.
03
Integrate according to the API documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limitations
Before formal use, understand the output quality and capability scope.
4k mode cannot be used together with camera_control. If the task focuses on explicitly controlling camera movements such as dolly moves and turns, choose std or pro; if the focus is 4K output, you must forgo this camera movement parameter and should not submit both configurations together.
Image-to-video requires start_image_url; end_image_url can only be used as a paired ending frame and cannot be used alone. First and last frames guide the start and end of the image, but do not specify the intermediate process frame by frame; complex changes should be broken down into more clearly defined shot tasks.
kling-v3 should not use Omni's multi-image reference or reference video editing workflow. image_list and video_list are creation methods for other models; enabling audio does not mean that voice-over, mixing, and subtitle production are complete, and sound-enabled videos still need to be reviewed and post-processed according to delivery requirements.
Frequently Asked Questions
Answers to common questions about using kling-v3.
How do I start generating videos with kling-v3?
Submit a request to POST /kling/videos and explicitly set model to kling-v3. For text creation, use action=text2video and fill in prompt; for image creation, use image2video and provide start_image_url, then choose duration, aspect ratio, and audio as needed.
What durations and aspect ratios can be generated?
Integer durations from 3–15 seconds are supported, along with three aspect ratios: 16:9, 9:16, and 1:1. You can compose shots for ads, vertical short videos, or square content; the duration should match the subject's actions, and too many continuous events should not be arranged within shorter clips.
How do I use start and end frames?
Select image2video and provide the starting frame start_image_url first, then add end_image_url when you need to specify the ending image. The end frame cannot be submitted without the start frame. The prompt can describe how the subject transitions from the initial state to the final state, keeping the visual guidance consistent with the action requirements.
Can 4K video control camera movement and generate sound at the same time?
4k mode supports generate_audio, but does not support camera_control. Explicitly enable audio generation when you need sound; use std or pro when you need parameterized camera movement. Audio is disabled by default, so selecting 4k alone will not automatically produce a video with sound.
How do I retrieve generated results, and can they be processed asynchronously?
You can set async=true to first obtain task_id, then query task progress and results; you can also configure callback_url to receive completion notifications. Completed results include information such as video_url, video_id, and state. Retrieve the video only after confirming success, rather than treating successful task submission as completed video generation.
Model information · Updated: 2026-10-01. See the API and pricing sections for invocation parameters and billing rules.