HomePhoto to Video › Create Image

How to generate an image from a text prompt

Describe what you want and the AI renders it as a still image at 1024 pixels, in any aspect ratio from square to 21:9. In the Photo to Video app for iPhone this is Create Image mode. You can also supply up to four reference images to steer the look — and because this app also animates stills, the result can go straight into a video.

An image generated from a written prompt
Generated from a written prompt. Nothing was uploaded.

Steps

  1. Open Create Image. In the app, go to Modes and choose Create Image.
  2. Describe the image. Say what is in it, and in what style.
  3. Optionally add references. Up to four images to guide subject or style.
  4. Pick a model and aspect ratio. GPT Image 2 or Flux; nine ratios from 1:1 to 21:9.
  5. Generate, then animate it if you want. The still can be fed into Photo to Video mode.

What it supports

InputA prompt, plus 0–4 optional reference images
OutputImage, 1024 pixels
Aspect ratios1:1, 9:16, 16:9, 4:3, 3:4, 3:2, 2:3, 21:9, 9:21
ModelsGPT Image 2 (default), Flux — Dev and Pro tiers
PromptRequired
PlatformiPhone (iOS). No Android or web version.
PricingCredits per image: 17 on GPT Image 2, 2 on Flux. No subscription.

Which model to use

The cost gap here is unusually wide: GPT Image 2 is 17 credits an image and follows a written description most literally, including text inside the image. Flux is 2 credits — roughly an eighth — and is fast and strong on style, which makes it the sensible place to iterate on wording before spending on a final render.

Getting a better result

  • Draft on Flux, finish on GPT Image 2. At 2 credits versus 17, there is no reason to iterate on the expensive model.
  • Pick the aspect ratio for where it will end up — 9:16 for a phone-screen video, 16:9 for anything landscape.
  • Name a medium and a light source: “oil painting, late afternoon side light” does more work than a list of adjectives.
  • If you plan to animate the result, leave some empty space around the subject — tightly cropped images have less room to move.

What people use it for

A still you intend to animate

Generating the frame first and animating it second gives you control over composition that one text-to-video pass does not. Leave space around the subject so there is room to move.

Prompt a lone red umbrella on an empty beach, overcast light, wide composition, 16:9

Cover art or a thumbnail at a specific shape

Nine aspect ratios, so you generate at the shape you need instead of cropping down to it.

Prompt bold graphic portrait, high contrast, single light source, 9:16

Iterating on a concept before committing

Flux at 2 credits against GPT Image 2 at 17 makes exploration nearly free. Settle the wording cheaply, render the final once.

Prompt concept sketch of a folding electric scooter, three-quarter view, studio grey

An image containing readable text

GPT Image 2 is the only model here that renders text with any reliability, and it earns its credits when the text is the point.

Prompt a vintage enamel sign reading OPEN ALL NIGHT, chipped paint, warm light

Where it struggles

Output is fixed at 1024 pixels — fine for screens, not for anything printed large. Flux is fast and cheap but follows instructions loosely; if the brief is precise, the cheap model will fight you. Neither model reliably reproduces a specific real person or brand.

Questions

Can this app generate images, not just video?

Yes. Create Image produces a still image at 1024 pixels from a prompt, with no video involved.

Which image model is cheaper?

Flux, at 2 credits per image versus 17 for GPT Image 2 — about an eighth of the cost.

Can I control the shape of the image?

Yes, nine aspect ratios from 1:1 to 21:9 and 9:21.

Can I use a reference image?

Yes, up to four, to guide the subject or the style.

Can a generated image then be animated?

Yes. Feed it into Photo to Video mode and the app animates it like any other still.