Home › Photo to Video › Fusion
How to combine several photos into one AI video
Select up to seven reference images — people, objects, backgrounds or style references — describe how they should relate, and the AI generates one scene containing them. In the Photo to Video app for iPhone this is Fusion mode. Four models support it and clips run 1 to 16 seconds.
Steps
- Open Fusion. In the app, go to Modes and choose Fusion.
- Add 1 to 7 references. Subjects, backgrounds, or images that carry the style you want.
- Describe the interaction. Say how the references relate — who is where, doing what, in which setting.
- Choose a model, length and resolution. Four models support this mode; see the table below.
- Generate and save. Save the composed clip to your camera roll.
What it supports
| Input | 1–7 reference images |
|---|---|
| Output | Video, 360p to 1080p depending on model |
| Length | 1–16 seconds depending on model |
| Models | Seedance 2.0 (default), Kling 3, Vidu Q3, Happy Horse |
| Prompt | Required — it describes how the references interact |
| Platform | iPhone (iOS). No Android or web version. |
| Pricing | Credits per second of output, by model. No subscription. |
Which model to use
Seedance 2.0 is the default and the most reliable at keeping several subjects recognisable at once. Vidu Q3 allows the longest clips. Kling 3 is the pick when the references are people and the result needs believable bodies. Happy Horse is the cheapest per attempt, useful while you are still working out the prompt.
Getting a better result
- Fewer references usually beats more. Two or three well-chosen images hold together better than seven competing ones.
- Say explicitly what each reference is for — “the woman from the first image, the beach from the second”. Left implicit, the model guesses.
- Use a clean cut-out or plain background for a subject you need kept intact.
- Reference images that disagree on lighting produce a composite that looks like one. Match them roughly.
What people use it for
Two people who were never photographed together
Supply each as a reference and say which is which. Seedance 2.0, the default, is the most reliable at keeping both recognisable.
Prompt the woman from the first image and the man from the second walk side by side along the shore
A product in a setting you do not have access to
One reference for the product, one for the location. Cheaper than a location shoot, and repeatable.
Prompt the sneaker from the first image sits on the wet rocks from the second, waves behind
Borrowing a look you cannot put into words
Use a reference purely for style and say so — otherwise the model treats it as a subject to include.
Prompt the subject from the first image, in the colour palette and grain of the second
Building a scene from parts
Subject, background and a prop as three references. Two or three beat seven — more references compete rather than combine.
Prompt the dog from the first image runs through the field in the second, ball from the third in its mouth
Where it struggles
Recognisability drops as you add references — seven is the ceiling, not the target. References that disagree on lighting produce a composite that looks composited. For a face that must stay exact, edit an image instead.
Questions
How many photos can Fusion combine?
Up to seven reference images in one generation.
Can it put two different people in the same video?
Yes. Supply each as a reference and describe how they interact. Keeping both recognisable is easiest on the default model, Seedance 2.0.
Can I use a photo just for its style?
Yes. A reference can be a subject, a background, or an image supplying the look, as long as your prompt says which role it plays.
How long can a Fusion clip be?
1 to 16 seconds depending on model; Vidu Q3 allows the longest.
Is a prompt required?
Yes. Fusion needs a prompt because it describes how the references combine.