HomePhoto to Video › Edit Image

How to edit a photo by describing the change

Upload a photo, write what should change, and the AI re-renders the image with that change applied — swap a background, alter clothing, restyle the scene — while keeping facial structure, lighting and colour consistent. In the Photo to Video app for iPhone this is Edit Image mode. It accepts up to four reference images and outputs a 1024-pixel image.

An AI-edited photograph where the scene was changed from a written instruction
Edited from a written instruction. Facial structure, lighting and colour were preserved.

Steps

  1. Open Edit Image. In the app, go to Modes and choose Edit Image.
  2. Add the photo to edit. Plus up to three more references to guide the result.
  3. Describe the change. Be specific about what should change — and, if it matters, what should not.
  4. Choose a model. GPT Image 2 by default, or Flux Pro.
  5. Generate and save. The result is an image, not a video.

What it supports

Input1–4 images
OutputImage, 1024 pixels
ModelsGPT Image 2 (default), Flux Pro
PromptRequired — it is the edit instruction
Aspect ratios1:1, 9:16, 16:9, 4:3, 3:4, 3:2, 2:3, 21:9, 9:21
PlatformiPhone (iOS). No Android or web version.
PricingCredits per image, by model. No subscription.

Which model to use

GPT Image 2 is the default: it follows written instructions most literally, which is what you want when the edit is specific. Flux Pro costs less than half as much per image and is quicker, making it the better choice for stylistic changes and for iterating on wording before committing.

Getting a better result

  • State what must stay as well as what changes — “keep her face and pose exactly, replace the background with a snowy street”.
  • One change at a time. Several instructions in one prompt tend to be applied partially.
  • Add a reference image for a look you cannot describe well in words.
  • Iterate on the cheaper model, then re-run the wording that worked on GPT Image 2.

What people use it for

Changing a background without reshooting

State that the subject must stay exactly as it is, or the model reinterprets them along with the background.

Prompt keep her face, pose and clothing exactly as they are; replace the background with a snowy street at dusk

Colour and material variants of a product

One photo becomes a range. Flux is cheap enough to generate every variant and keep what works.

Prompt change the chair upholstery to olive green linen, keep the frame, lighting and shadows identical

Removing something you cannot crop out

A sign, a bystander, a cable. Say what should be there instead, not only what to remove.

Prompt remove the parked car; continue the brick wall and pavement behind it

Getting a still ready to animate

Fix the frame first, then send it to Photo to Video. Cleaning up a still is far cheaper than regenerating a video that inherited the problem.

Prompt remove the date stamp in the corner, keep everything else untouched

Where it struggles

Several instructions in one prompt get applied partially — one change per pass. Output is 1024 pixels, so this is not a retouching tool for print. Text in the image is as unreliable here as anywhere else.

Questions

Can AI edit a photo from a text instruction?

Yes. Edit Image re-renders the photo with your described change applied, preserving facial structure, lighting and colour tone.

Does it output a video?

No. This mode produces an image at 1024 pixels. The video modes are separate.

How many reference images can I use?

Up to four, including the photo being edited.

Which model edits photos best?

GPT Image 2, the default, follows instructions most literally. Flux Pro is cheaper and faster, better for style changes and for testing prompts.

Will it change the person's face?

It is designed not to — facial structure is preserved unless your prompt asks otherwise.