HomePhoto to Video › Fusion

How to combine several photos into one AI video

Select up to seven reference images — people, objects, backgrounds or style references — describe how they should relate, and the AI generates one scene containing them. In the Photo to Video app for iPhone this is Fusion mode. Four models support it and clips run 1 to 16 seconds.

A frame from an AI-generated video that combined several separate reference images into one scene
Several reference images, one generated scene.

Steps

  1. Open Fusion. In the app, go to Modes and choose Fusion.
  2. Add 1 to 7 references. Subjects, backgrounds, or images that carry the style you want.
  3. Describe the interaction. Say how the references relate — who is where, doing what, in which setting.
  4. Choose a model, length and resolution. Four models support this mode; see the table below.
  5. Generate and save. Save the composed clip to your camera roll.

What it supports

Input1–7 reference images
OutputVideo, 360p to 1080p depending on model
Length1–16 seconds depending on model
ModelsSeedance 2.0 (default), Kling 3, Vidu Q3, Happy Horse
PromptRequired — it describes how the references interact
PlatformiPhone (iOS). No Android or web version.
PricingCredits per second of output, by model. No subscription.

Which model to use

Seedance 2.0 is the default and the most reliable at keeping several subjects recognisable at once. Vidu Q3 allows the longest clips. Kling 3 is the pick when the references are people and the result needs believable bodies. Happy Horse is the cheapest per attempt, useful while you are still working out the prompt.

Getting a better result

  • Fewer references usually beats more. Two or three well-chosen images hold together better than seven competing ones.
  • Say explicitly what each reference is for — “the woman from the first image, the beach from the second”. Left implicit, the model guesses.
  • Use a clean cut-out or plain background for a subject you need kept intact.
  • Reference images that disagree on lighting produce a composite that looks like one. Match them roughly.

What people use it for

Two people who were never photographed together

Supply each as a reference and say which is which. Seedance 2.0, the default, is the most reliable at keeping both recognisable.

Prompt the woman from the first image and the man from the second walk side by side along the shore

A product in a setting you do not have access to

One reference for the product, one for the location. Cheaper than a location shoot, and repeatable.

Prompt the sneaker from the first image sits on the wet rocks from the second, waves behind

Borrowing a look you cannot put into words

Use a reference purely for style and say so — otherwise the model treats it as a subject to include.

Prompt the subject from the first image, in the colour palette and grain of the second

Building a scene from parts

Subject, background and a prop as three references. Two or three beat seven — more references compete rather than combine.

Prompt the dog from the first image runs through the field in the second, ball from the third in its mouth

Where it struggles

Recognisability drops as you add references — seven is the ceiling, not the target. References that disagree on lighting produce a composite that looks composited. For a face that must stay exact, edit an image instead.

Questions

How many photos can Fusion combine?

Up to seven reference images in one generation.

Can it put two different people in the same video?

Yes. Supply each as a reference and describe how they interact. Keeping both recognisable is easiest on the default model, Seedance 2.0.

Can I use a photo just for its style?

Yes. A reference can be a subject, a background, or an image supplying the look, as long as your prompt says which role it plays.

How long can a Fusion clip be?

1 to 16 seconds depending on model; Vidu Q3 allows the longest.

Is a prompt required?

Yes. Fusion needs a prompt because it describes how the references combine.