Home › Photo to Video › Edit Image
How to edit a photo by describing the change
Upload a photo, write what should change, and the AI re-renders the image with that change applied — swap a background, alter clothing, restyle the scene — while keeping facial structure, lighting and colour consistent. In the Photo to Video app for iPhone this is Edit Image mode. It accepts up to four reference images and outputs a 1024-pixel image.
Steps
- Open Edit Image. In the app, go to Modes and choose Edit Image.
- Add the photo to edit. Plus up to three more references to guide the result.
- Describe the change. Be specific about what should change — and, if it matters, what should not.
- Choose a model. GPT Image 2 by default, or Flux Pro.
- Generate and save. The result is an image, not a video.
What it supports
| Input | 1–4 images |
|---|---|
| Output | Image, 1024 pixels |
| Models | GPT Image 2 (default), Flux Pro |
| Prompt | Required — it is the edit instruction |
| Aspect ratios | 1:1, 9:16, 16:9, 4:3, 3:4, 3:2, 2:3, 21:9, 9:21 |
| Platform | iPhone (iOS). No Android or web version. |
| Pricing | Credits per image, by model. No subscription. |
Which model to use
GPT Image 2 is the default: it follows written instructions most literally, which is what you want when the edit is specific. Flux Pro costs less than half as much per image and is quicker, making it the better choice for stylistic changes and for iterating on wording before committing.
Getting a better result
- State what must stay as well as what changes — “keep her face and pose exactly, replace the background with a snowy street”.
- One change at a time. Several instructions in one prompt tend to be applied partially.
- Add a reference image for a look you cannot describe well in words.
- Iterate on the cheaper model, then re-run the wording that worked on GPT Image 2.
What people use it for
Changing a background without reshooting
State that the subject must stay exactly as it is, or the model reinterprets them along with the background.
Prompt keep her face, pose and clothing exactly as they are; replace the background with a snowy street at dusk
Colour and material variants of a product
One photo becomes a range. Flux is cheap enough to generate every variant and keep what works.
Prompt change the chair upholstery to olive green linen, keep the frame, lighting and shadows identical
Removing something you cannot crop out
A sign, a bystander, a cable. Say what should be there instead, not only what to remove.
Prompt remove the parked car; continue the brick wall and pavement behind it
Getting a still ready to animate
Fix the frame first, then send it to Photo to Video. Cleaning up a still is far cheaper than regenerating a video that inherited the problem.
Prompt remove the date stamp in the corner, keep everything else untouched
Where it struggles
Several instructions in one prompt get applied partially — one change per pass. Output is 1024 pixels, so this is not a retouching tool for print. Text in the image is as unreliable here as anywhere else.
Questions
Can AI edit a photo from a text instruction?
Yes. Edit Image re-renders the photo with your described change applied, preserving facial structure, lighting and colour tone.
Does it output a video?
No. This mode produces an image at 1024 pixels. The video modes are separate.
How many reference images can I use?
Up to four, including the photo being edited.
Which model edits photos best?
GPT Image 2, the default, follows instructions most literally. Flux Pro is cheaper and faster, better for style changes and for testing prompts.
Will it change the person's face?
It is designed not to — facial structure is preserved unless your prompt asks otherwise.