A short video model that consistently renders subjects and scenes using multiple image references
HappyHorse-1.0-R2V is a video generation model for reference-image-based creation. It combines the subjects and scenes in images with action intent described in text to generate new clips. It is suited for creative tasks with existing character images, product images, or visual concepts, focusing on preserving reference relationships rather than simply animating a single image. On this platform, you can explicitly select this version to complete multi-image input, aspect ratio settings, and video result retrieval.
Clarify capacity, input/output, and invocation methods before choosing a model.
Creation mode
Reference-image-to-video; image and text input, video output
Number of reference images
1–9 images, submitted via image_urls
Output resolution
720P, 1080P
Platform generation duration
3–15 seconds
Platform aspect ratios
16:9, 9:16, 1:1, 4:3, 3:4
Reference identifiers
Names such as character1 and character2 can be used according to image order
Task delivery
Supports asynchronous queries, completion callbacks, and video URL retrieval
Multi-image reference and image-text generation are model capabilities; duration, aspect ratio, and task delivery methods are listed according to this platform's supported invocation scope.
Core Capabilities
Learn what happyhorse-1.0-r2v can bring to your work.
Ground subjects and scenes in visual references
Creation does not have to rely solely on text descriptions of appearance. Submit subject images, scene images, and action prompts together, and the model can generate clips based on clear visual references, with a focus on the consistency of subject and scene references. It is suitable for short videos that need to preserve character recognizability, product form, or visual direction, but it does not mean pixel-perfect image copying.
Use multiple images to organize creative intent
A single task can include 1–9 reference images, and prompts can refer to them in order using names such as character1 and character2. You can specify which image a subject comes from, which scene or visual elements to use, and then describe the action and camera work, reducing ambiguity in creative relationships when using multiple image inputs.
Plan generation around short-video delivery
The platform offers generation durations of 3–15 seconds, 720P and 1080P resolutions, and landscape, portrait, and square aspect ratios. Tasks can be submitted asynchronously, with results retrieved through queries or callbacks, making it suitable for integrating reference-image creation into asset management, content review, and download workflows without having to wait continuously for requests to finish.
Use Cases
Start with specific tasks to find where the model can be effective.
Product showcase clips
Provide product reference images and images of the display environment, then use text to describe product placement, camera movement, and visual atmosphere to generate short videos suitable for marketing proposal previews. Reference images convey appearance, while prompts organize actions; before delivery, carefully check logos, structure, and small text, and avoid treating generated visuals directly as precise product records.
Character storyboard previsualization
Combine character design images with scene reference images, and describe single-sequence actions such as a character entering the frame, turning around, or walking to generate storyboard previsualization material. This is suitable for discussing the relationship between characters and environments before filming or animation production; multiple shots can be generated separately, then checked for continuity in character appearance rather than assuming continuity is automatically maintained across tasks.
Multi-format assets for the same theme
Using the same set of reference images, organize landscape, portrait, or square compositions separately to generate short-video candidates for different display placements. Each prompt should specify the subject position and camera focus to make aspect-ratio tradeoffs easier to compare. Output videos can enter the editing workflow for further additions such as subtitles, sound, and brand information.
How to choose this model
Choose based on task complexity, input materials, and expected results.
Multiple image references or a fixed first frame
When you need to combine character, prop, and scene references, choose 1.0-r2v; if the key requirement is to make a certain image the first frame of the video, choose an i2v variant. If you only have a text concept, consider t2v; if you already have a video and want to change outfits or modify the visual style, choose video-edit. They correspond to different input methods, and cannot be considered the same capability simply by changing the action name.
Explicitly choose 1.0 instead of relying on the default version
1.0-r2v and 1.1-r2v are different model variants, and reference-image generation defaults to 1.1. If you already have prompts and material test sets built around 1.0, you can explicitly specify 1.0 to continue evaluation; when trying 1.1, it is recommended to use the same reference images and tasks to compare subject retention and camera expression, rather than directly interpreting version numbers as a fixed degree of quality improvement.
Get started
From a small-scale task to formal integration.
01
Prepare tasks and materials
Define the objective, required inputs, and output requirements, using real business examples as a starting point.
02
Try it in the API testing area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to view the results.
03
Integrate according to the API documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage limitations
Understand output quality and capability scope before formal use.
Reference images are used to guide generation and do not mean that subjects, clothing, logos, or scene details will remain completely unchanged. If multiple images conflict with each other in appearance or style, first filter the materials and explain the purpose of each image in the prompt; content with high requirements for product details and character recognizability needs to be checked segment by segment.
This model is for reference-image-to-video generation, not an existing video editing entry point or a fixed-first-frame animation entry point. Do not treat video_url or original-audio retention settings as its primary creation method; when you need to modify existing clips or retain the original audio, choose the corresponding video editing model.
Platform-generated clips are 3–15 seconds long; long narratives need to be split into multiple tasks and edited afterward. Image URLs must be publicly accessible; local paths or links that require login are not suitable as material inputs. After submission, you must also distinguish between pending, successful, and error states, and cannot judge that a video has been generated based only on the task ID.
Frequently Asked Questions
Answers to common questions when using happyhorse-1.0-r2v.
What must be specified when calling 1.0-r2v?
Submit a request to /happyhorse/videos, explicitly set model to happyhorse-1.0-r2v, action to reference_to_video, and provide a prompt and 1–9 image_urls. Do not rely on default actions or models, as they may not match reference-image generation tasks.
Can it be used with only one reference image?
Yes, the number of reference images starts at 1. One image is suitable for providing a subject or visual direction, while multiple images are suitable for supplementing different reference relationships. If you require the image to directly become the first frame, choose the i2v variant; R2V focuses on using reference information to create new clips.
How can prompts accurately refer to multiple images?
Following the order in image_urls, use names such as character1 and character2 to refer to the corresponding images, and clearly specify the purpose of each. For example, state that the subject references the first image and visual elements reference the second image, then describe the action and camera. After adjusting the image order, also update the prompt accordingly.
Can it directly generate dialogue and preserve the original video audio?
Do not design tasks assuming that dialogue, lip-sync, or original audio preservation are confirmed features of 1.0-r2v. Its stated purpose is generating video from image and text references; preserving existing video audio is an operation of video-edit. When specific audio content is needed, dubbing and mixing can be completed in post-production.
How do I get the video after asynchronous submission?
After submitting with async, retain the task_id and query it through /happyhorse/tasks; you can also provide callback_url to receive the completed result. Task statuses include pending, succeeded, and error. After success, obtain video_url and use the returned duration and resolution to check whether the final video meets this task's requirements.
Model information · Updated: 2026-10-01. For calling parameters and billing rules, see the API and pricing sections.