Generate previewable video assets from text-based ideas
minimax-t2v is MiniMax's text-to-video entry point, designed for creative tasks without first-frame images that aim to turn text concepts directly into dynamic visuals. Enter descriptions of the subject, environment, actions, and visual atmosphere to start video generation and obtain a finished video link. It is suitable for concept previews, creative proposals, and asset exploration; compared with reference-image-driven minimax-i2v, the creative starting point is text rather than existing images.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Text-to-video, no first-frame reference image required
Content input
prompt text prompt
Generation request
POST /hailuo/videos;model=minimax-t2v;action=generate
Task execution
Supports async=true, returns task_id first
Completion notification
Supports callback_url for receiving POST JSON results
Result delivery
video_url video link and task status information
This lists the actual creation and request methods on this platform and does not treat the native parameters of other Hailuo versions as specifications for this model.
Core capabilities
No need to create an image first—start with a description
minimax-t2v uses text as the creative starting point, making it suitable for the ideation stage when image assets are not yet available. You can structure descriptions around “who does what in which environment,” then add lighting, color tone, and visual atmosphere. Turn ideas into viewable dynamic drafts first, then decide whether to continue creating reference images or proceed with subsequent editing.
Iterate on visuals through prompts
Creative input is concentrated in the prompt, making it easy to retain the same theme while progressively adjusting scene or action descriptions. It is recommended to specify one main change each time, such as changing a sunset coast to an overcast coast, then comparing the results. This iterative approach is suitable for exploring visual directions, but it is not equivalent to precisely locking camera movement or controlling frame by frame.
Keep generation and result management separate
An asynchronous task workflow is supported: save the task_id after submission, then query the status or receive a completion callback. The video_url in the result is used to obtain the video, while the task identifier corresponds to the original idea. For applications that need to submit multiple pieces of copy continuously, this approach makes it convenient to separate generation from preview, review, and asset archiving.
Use Cases
Dynamic Pitches for Ad Concepts
Rewrite a visual segment from an ad script into descriptions of the subject, scene, and action, such as a person walking down a street on a rainy night, then generate a video draft for team discussion. The focus of the deliverable is to present the creative direction and atmosphere, rather than directly replace a final ad production that requires accurate trademarks, packaging text, and character identities.
Concept Previsualization for Story Scenes
Enter scene descriptions from a story into the model to explore how environments, actions, and visual atmosphere can be combined. Suitable for creating dynamic references for individual plot points before live-action filming or animation production. If the story requires the same character to appear consistently, it is recommended to first use this model to explore directions, then plan subsequent shots using a reference-image-based creation approach.
Asset Exploration in Content Production
Write prompts around ocean waves, urban street scenes, or other themes, generate candidate videos, and preview them via links to select assets suitable for subsequent editing. Different prompts and task results can be archived together for easy comparison and reuse. Subtitles, music, voice-over, and complete timeline editing should be planned as separate production steps.
How to Choose This Model
When You Have No Images, Start with Text
If you only have a creative idea or a scene description, minimax-t2v better matches your current input conditions, so you do not need to prepare a first frame before submitting a task. If you already have a product image, character image, or a defined composition and want to use it as the video starting point, minimax-i2v is more suitable. The main difference between the two is the creative basis; text-to-video should not be treated as animation processing for a specified image.
Choose Another Creation Method for Fine Control Needs
minimax-t2v is suitable for text-driven visual exploration and should not be used as Director mode. If the task starts from a reference image and emphasizes enhanced creative control, consider minimax-i2v-director. When choosing, first consider whether image constraints are needed, then consider control requirements; do not assume that different models have the same camera capabilities simply because they belong to the Hailuo series.
Get Started
Prepare the Creative Starting Point for This Model
Start with a text description of the subject, action, and shot brief; no first-frame image is required.
Use the Hailuo Video Endpoint
Specify /hailuo/videos with model=minimax-t2v, action=generate, and a prompt; this is a text-generation task, so it does not require a first-frame URL and does not use the MiniMax H3 content structure.
Retrieve the Video Link
Set async=true to first obtain the task_id, query it through /hailuo/tasks or receive results via callback_url; read the successful video_url, then proceed to editing after a complete review.
Trial suggestion: Concept shots without reference images
Input and goal
A lighthouse by the sea at dusk gradually lights up, the camera slowly pushes forward, the sea is calm, keep to one scene, and do not add subtitles.
Acceptance and next steps
Start with action=generate and a text prompt; when there are no model-specific duration/resolution fields, do not copy H3 parameters.
Usage boundaries
Text descriptions are not a strict lock on a person's identity, product appearance, or composition. If the task requires faithfully continuing a particular image, use the reference-image creation entry point; this model is better suited to exploring visual approaches first and then manually selecting results, rather than handling the precise reproduction of a specified image.
Do not treat camera descriptions in prompts as dedicated camera-control instructions. Complex continuous camera movements, multiple actions occurring simultaneously, or strict frame-by-frame arrangements should be split into clearer independent ideas for testing; do not expect this model to have the control capabilities of the Director model.
The core deliverable is the generated video and its result link; voice-over, sound effects, lip-sync, or a complete editing timeline should not be treated as default deliverables. When these production steps are needed, arrange separate audio and editing workflows, and check before use whether the actual video meets project requirements.
Frequently Asked Questions
Does minimax-t2v require uploading an image?
No, it generates videos from text prompts. It is recommended to first clearly describe the subject, environment, and main actions, then add atmosphere descriptions. If your task must start from an existing image, choose minimax-i2v instead of treating a first-frame image as a required input for minimax-t2v.
What is the difference between it and minimax-i2v?
minimax-t2v starts from a text concept; minimax-i2v starts from a reference image and requires an image link. Choose the former when you have no visual assets and want to explore ideas; choose the latter when you already have a specific image and want to create motion around it, as it better matches the task requirements.
Is minimax-t2v a Director version?
It should be used through the text-to-video entry point and should not be regarded as a Director model. For reference-image-driven Director creation, you can choose minimax-i2v-director. Even if you include camera movement requirements in the prompt, that does not mean you gain Director-exclusive preset shots or fine-grained motion control.
How do I write prompts suitable for it?
It is recommended to organize the prompt around one clear scene, describing the main subject, the action taking place, the environment, and the visual atmosphere. First avoid mutually contradictory action requirements, then compare results by changing one element at a time. This approach makes iteration easier; it does not mean complex instructions can necessarily be executed precisely.
How do I retrieve the generated video after submission?
When calling, you can set async=true, first record the returned task_id, then query the task result; you can also configure callback_url to receive completion notifications. After obtaining the result, use the status information to determine whether the task succeeded, and retrieve the video through video_url in data; do not treat the task ID as the finished video URL.