HomePhoto to Video › Text to Video

How to make a video from a text prompt

Write a description of the scene you want and the AI generates the video from nothing else — no photo, no footage. In the Photo to Video app for iPhone this is Text to Video mode. Nine models are available, clips run 1 to 16 seconds, and Veo 3.1 can render 4K with sound.

A frame from a video generated entirely from a written description
No photo was used. This was generated from a sentence.

Steps

  1. Open Text to Video. In the app, go to Modes and choose Text to Video.
  2. Describe the scene. Say what is in frame, what moves, and how the camera behaves.
  3. Choose a model, length and resolution. Nine models support this mode; see the table below.
  4. Generate. The clip is built from the prompt alone.
  5. Save it. Save the result to your camera roll.

What it supports

InputA text prompt only
OutputVideo, 480p to 4K depending on model
Length1–16 seconds depending on model
ModelsSeedance 1.0 Pro (default), Seedance 2.0, Veo 3.1, Kling 3, Wan 2.7, Vidu Q3, Hailuo-02, Grok Imagine, Happy Horse
PromptRequired
PlatformiPhone (iOS). No Android or web version.
PricingCredits per second of output, by model. No subscription.

Which model to use

Seedance 1.0 Pro is the default and cheapest per second. Veo 3.1 is the choice for 4K and for generated audio. Seedance 2.0 holds up best on prompts with several things happening at once. Kling 3 is strong on human figures. For quick drafts of an idea, Hailuo-02 costs the least per attempt.

Getting a better result

  • Write the shot, not the story. Subject, action, setting, camera — “low-angle shot of a red kite rising over wet sand, camera tilts up” beats a paragraph of narrative.
  • One action per clip. Prompts asking for a sequence of events tend to produce neither.
  • Name the look explicitly — “shot on 35mm, shallow depth of field”, “stop-motion”, “drone footage”.
  • Iterate short. Generate at 4 seconds until the prompt is right, then re-run it longer.

What people use it for

A shot you need but have no footage for

An establishing shot, a texture, a mood insert. Faster than sourcing stock and specific to your brief.

Prompt low-angle shot of a red kite rising over wet sand at dusk, camera tilts up, 35mm

B-roll to cut between talking-head footage

Generate several 4-second variations and keep the two that cut cleanly. This is where the cheaper models earn their place.

Prompt hands close a laptop on a wooden desk, warm window light, shallow depth of field

Testing an ad idea before committing to a shoot

A rough generated version answers whether the concept reads at all, for a fraction of a day's production.

Prompt a runner crosses an empty bridge at sunrise, camera tracks alongside, breath visible

A clip that needs its own sound

Veo 3.1 generates synchronised audio in the same pass, which no other model here does.

Prompt rain on a tin roof at night, a single lamp swinging, distant thunder

Where it struggles

Text inside the frame comes out wrong more often than right. Prompts asking for a sequence of events tend to deliver neither half. Anything needing a specific real person, logo or place will be an approximation, not that thing.

Questions

Can AI make a video from just text?

Yes. Text to Video generates the footage from a written description alone, with no photo or video input.

How long can a text-to-video clip be?

1 to 16 seconds depending on the model. Veo 3.1 is fixed at 4, 6 or 8 seconds.

Can the video have sound?

Yes, with Veo 3.1, which generates synchronised audio. Any clip can also have a soundtrack added afterwards with Add Audio.

Which model is best for text to video?

Seedance 1.0 Pro for cost, Veo 3.1 for 4K and audio, Seedance 2.0 for complex scenes. You can switch per generation at no extra cost.

Does it work on Android?

No. The app is iPhone-only.