A Kling Creative Model Combining Multi-Image References and Video Editing
Kling O1 is Kuaishou Kling's multimodal video model for generation and editing, suitable for combining images of people, products, and scenes with existing clips into clear creative instructions. Its value goes beyond animating static images, including adjusting video style, visual elements, and composition with reference materials. This platform's kling-o1 entry provides 5-second short video creation, ideal for refining ads and serial content shot by shot.
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Creation methods
Text-to-video, image-to-video, multi-image references, reference video, and video editing
Generation duration and modes
Platform entry: 5 seconds; std / pro
Aspect ratios
16:9, 9:16, 1:1
Image references
Up to 7 images without a reference video; up to 4 images with a reference video
Video reference
Up to 1 MP4/MOV clip; 3–10 seconds; ≤200MB; 24–60fps; width and height each 700–2160px
Audio handling
Does not generate audio; original audio can be retained or removed for video references
Result delivery
Video link, video ID, task ID, and task status; supports asynchronous queries and callbacks
O1's multimodal generation and editing positioning corresponds to its creative capabilities; the durations, asset requirements, and modes above are the invocation specifications for this platform entry.
Core Capabilities
Learn what kling-o1 can bring to your work.
Bring characters, scenes, and styles into the same shot
Multiple image references can separately provide character appearance, product form, scenes, and visual style, while text can describe their relationship within the shot. Compared with using only long prompts to describe appearances, this approach is better suited to creating a series of short videos around the same character or product; reference materials still need to be explicitly cited by their corresponding numbers in the prompt.
Make targeted edits around existing clips
Set a reference video as the base to add, remove, or modify elements around existing footage, adjust composition, and change style, color, and weather. Instructions should explain both what you want to change and what you want to preserve—for example, changing the overall visual style while preserving the original motion and composition—making the editing goal clearer than a vague request to “optimize the video.”
Use video characteristics to guide new shots
Set a video as a feature to reference its style, camera movement, or direction for the next shot, rather than treating it directly as a clip to be modified. This is suitable for continuing creation based on an existing visual concept. Image-to-video can also specify first and last frames, providing anchors for the starting and ending images, but this cannot be combined with base video editing.
Use Cases
Start with specific tasks to find where the model can be effective.
Short-shot variations for product advertising
Input product reference images, set images, and clear action descriptions to create short shots suitable for landscape, portrait, or square placements. You can keep the product and set references unchanged while trying different actions or moods, then use the resulting videos for ad editing; packaging text and small labels should be checked frame by frame before delivery.
Visual revisions for existing footage
Input a short video that meets the requirements, set it as the base, and specify the colors, weather, composition, or style that need to change, as well as the subjects and motion that should be preserved. The deliverable is a regenerated video clip, suitable for comparing creative approaches and revising visuals; when the original sound needs to be retained, you can separately choose to preserve the original audio.
Shot continuity for serialized content
Prepare consistent character and scene reference images, or use the previous video as a feature reference, then describe the action and visual direction of the next shot. Produce one short video at a time, then organize them into serialized content during editing. Consistent reference materials and instruction structure help maintain recognizability, but continuity between shots still needs to be checked.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Prioritize O1 When You Have Reference Materials
If the core task is to combine references for people, products, and scenes, or to make targeted edits to existing clips, O1's multi-image and video reference workflow is worth trying first. If you only have a first frame, you can use image2video; for combining multiple materials and video editing, use text2video with a reference list—there is no need to classify every material-based task as image-to-video.
Choose V3 Omni When You Need Longer Clips or Audio
kling-o1 and kling-v3-omni are independent models, not mode switches for the same model. O1 is suitable for 5-second reference-based creation and editing; if you need 3–15-second generation, synchronized audio, or native 4K, consider the relevant V3 Omni workflow. Your choice should be based on material conditions and delivery requirements, rather than version names alone.
Get Started
From a small-scale task to formal integration.
01
Prepare the Task and Materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Debugging Area
Open the trial page, confirm the parameters supported by this endpoint, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limitations
Before formal use, understand the output quality and capability scope.
The O1 endpoint generates only 5-second clips and does not support 4k mode, audio generation, or camera_control. Video feature references can guide motion direction, but they are not equivalent to numerical camera movement control; longer works should be organized through storyboard generation and post-production editing.
Omni-reference requests cannot use negative_prompt, cfg_scale, or camera_control. Reference images and videos must also be cited by sequence number in the prompt; simply uploading materials does not apply them automatically, and mixing first/last-frame fields with image lists may also change the material order.
Reference videos must meet format, duration, frame rate, and size requirements. Base editing cannot additionally specify first and last frames; the final frame for image-to-video must be used together with a first frame. Images must be JPG/JPEG/PNG, ≤10MB, with a shortest side ≥300px; you cannot directly submit arbitrary materials.
Frequently Asked Questions
Answers to common questions about using kling-o1.
Which action does Kling O1 use for video editing?
For reference video editing, use action=text2video, explicitly specify model=kling-o1, and set the asset as base in video_list. The prompt must reference that asset and describe the intended modifications; feature is used for feature reference, not direct editing of the base clip.
Why didn't my reference image take effect?
Assets in the reference list need to be referenced in the prompt by their corresponding index; simply passing image_list or video_list will not apply them automatically. You should also avoid mixing separate first/last frame fields with image lists, as this may change the asset order and cause instruction references to differ from expectations.
Can O1 generate audio or preserve the original video audio?
O1 does not support synchronized audio generation, so generate_audio should remain false. When using a reference video, you can choose to retain or remove the original audio through keep_original_sound; this is processing of the source asset's audio and does not mean the model will generate new dialogue, music, or sound effects.
If the reference video is up to 10 seconds, can the output also be 10 seconds?
You cannot treat the reference asset duration as the output duration. O1 accepts reference videos from 3–10 seconds, but the generation duration for this endpoint is 5 seconds. When preparing assets, also check the video dimensions, frame rate, format, and file size to avoid failure at submission due to unmet asset requirements.
How do I retrieve generated results in a business system?
Submit a task to POST /kling/videos and explicitly specify kling-o1. Setting async=true lets you obtain task_id first and then query the task; you can also use callback_url to receive the completed result. Check state and success to determine whether it is complete, then use video_url to retrieve the video.
Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.