A short video model that turns static images into controllable creative starting points
Grok Imagine Video 1.5 is xAI's video generation model, and grok-imagine-video-1.5:official focuses on creating short videos starting from images. Upload product photos, character images, or storyboard frames, then describe the action and camera movement in text to generate 1–15 second videos with support for up to 1080p. It is suitable for animating and creatively validating existing visual assets.
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Creation method
Image-to-video; image_url is required
Motion guidance
Optional prompt describing motion, camera work, and scene changes
Video duration
1–15 seconds, 6 seconds by default
Output resolution
480p, 720p, 1080p; 480p by default
Task delivery
Supports asynchronous submission, callbacks, and task queries
Invocation and results
POST /grok/videos; returns video_url upon completion
The above are the input and output specifications for this platform's invocation endpoint and do not equate the series' native capabilities with the features available through this endpoint.
Core capabilities
Learn what grok-imagine-video-1.5:official can bring to your work.
Design motion from existing images
Use an image to establish the creative starting point, then add subject movement, camera motion, and environmental changes with prompts. Compared with reimagining a scene from text, this approach is better suited to projects with existing product photos, character designs, or storyboard materials, placing the creative focus on how static images become dynamic shots.
Validate short video results in stages
Offers different resolution and short video duration options, making it easier to validate motion direction before producing delivery candidates. You can first use the default 480p and 6 seconds to check whether the shot matches your concept, then adjust the image or motion description and select a higher resolution, rather than investing in full production from the start.
Integrate generation into task workflows
Generation tasks can be handled asynchronously, without keeping the request connection open. Save the task_id after submission, receive results through task queries or callbacks, and then obtain the video link based on the status. This approach is suitable for asset management, batch creative testing, and applications that need to record the result of every generation.
Applicable Scenarios
Start with specific tasks to find where the model can be effective.
Dynamic Product Photo Display
Input an already captured product image, describe slow movement, background changes, or display actions, and generate short video candidates for review. First check the product outline, branding, and camera movement, then decide whether to use it in advertising edits. This is suitable for extending static visual assets into dynamic creative content.
Motion Previsualization for Storyboard Frames
Use a single storyboard frame or concept image as input, add descriptions such as a character turning, the camera moving closer, or environmental motion, and obtain playable shot drafts. The key deliverable is to discuss action and pacing rather than replace final filming; teams can use this to judge whether a shot design is worth developing further.
Visual Variants for Social Content
Based on the same key visual, try different action and camera descriptions to create multiple short video candidates. Input images should account in advance for the target aspect ratio and subject placement. After generation, select clips suitable for distribution, then complete subtitles, music, and editing to create publishable content assets.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Consider It First When You Have Image Assets
If a project already has clear visual imagery and you want to create short videos up to 1080p around it, you can choose 1.5:official. If you only have a text concept and need direct text-to-video generation, consider grok-imagine-video:official; it supports both text-to-video and image-to-video, and also covers tasks from 1–15 seconds.
Balance Input Method and Duration
For image-to-video short clips of 15 seconds or less, you can use this model; for longer clips or text-to-video work at the same time, consider grok-imagine-video-1.5-fast:reverse, which has a duration range of 6–30 seconds. Choose between them based on available materials, target length, and actual visual results, rather than treating version names as a quality ranking for all tasks.
Get Started
From a small-scale task to formal integration.
01
Prepare the Task and Materials
Clarify the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Debugging Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the full model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limits
Understand the output quality and capability scope before formal use.
This model must be provided with image_url; it cannot begin generation from text alone. reference_image_urls cannot replace the base input image; if you only have text-based ideas, choose a model that supports text-to-video, or first prepare an image suitable as a starting point for creation.
Each generation should be kept within 1–15 seconds; do not directly apply the 30-second capability of other models. Longer narratives are best split into separate shots, generated and edited individually; continuity of characters, composition, and actions across shots still requires manual review.
Aspect ratio design should start with the input image, with cropping, subject placement, and background space handled in advance. Higher output resolution does not guarantee accurate product details or character movements; when text, logos, and key demonstration actions are involved, review each segment before using it for delivery.
Frequently Asked Questions
Answers to common questions when using grok-imagine-video-1.5:official.
Can 1.5:official generate videos using text only?
No. This model uses image-to-video and requires image_url. Text prompts are used to supplement the actions, camera work, or scene changes in the image. If you do not yet have image assets, use grok-imagine-video:official or fast:reverse, which support text-to-video.
Do I still need to write a prompt in addition to the image?
Image-to-video can be used without a prompt, but adding one is recommended when you have clear motion goals. Descriptions should focus on how the subject moves, how the camera changes, and what happens in the scene. Avoid requesting too many conflicting actions at once, so the goal of the short video is easier to assess.
How long and at what resolution can videos be generated?
This model supports 1–15 seconds, with 6 seconds by default; available resolutions are 480p, 720p, or 1080p, with 480p by default. It is suitable to first create short drafts to check movement, then generate higher-resolution candidates for satisfactory ideas. The specific footage still needs to be reviewed in the end.
How do I retrieve asynchronously generated videos?
After setting async to true, save the returned task_id, then query the results through /grok/tasks; you can also provide callback_url to receive completion notifications. Only after the status reaches succeeded and video_url has been obtained should the result be considered a downloadable video asset.
How do I ensure I am calling this model?
When submitting a request to /grok/videos, explicitly set model to grok-imagine-video-1.5:official and provide image_url at the same time. Do not omit the model or use a name containing preview; retain task_id and trace_id to help link assets, task status, and generation results.
Model information · Updated: 2026-10-01. For request parameters and billing rules, see the API and pricing sections.