Home › Photo to Video › Text to Video
How to make a video from a text prompt
Write a description of the scene you want and the AI generates the video from nothing else — no photo, no footage. In the Photo to Video app for iPhone this is Text to Video mode. Nine models are available, clips run 1 to 16 seconds, and Veo 3.1 can render 4K with sound.
Steps
- Open Text to Video. In the app, go to Modes and choose Text to Video.
- Describe the scene. Say what is in frame, what moves, and how the camera behaves.
- Choose a model, length and resolution. Nine models support this mode; see the table below.
- Generate. The clip is built from the prompt alone.
- Save it. Save the result to your camera roll.
What it supports
| Input | A text prompt only |
|---|---|
| Output | Video, 480p to 4K depending on model |
| Length | 1–16 seconds depending on model |
| Models | Seedance 1.0 Pro (default), Seedance 2.0, Veo 3.1, Kling 3, Wan 2.7, Vidu Q3, Hailuo-02, Grok Imagine, Happy Horse |
| Prompt | Required |
| Platform | iPhone (iOS). No Android or web version. |
| Pricing | Credits per second of output, by model. No subscription. |
Which model to use
Seedance 1.0 Pro is the default and cheapest per second. Veo 3.1 is the choice for 4K and for generated audio. Seedance 2.0 holds up best on prompts with several things happening at once. Kling 3 is strong on human figures. For quick drafts of an idea, Hailuo-02 costs the least per attempt.
Getting a better result
- Write the shot, not the story. Subject, action, setting, camera — “low-angle shot of a red kite rising over wet sand, camera tilts up” beats a paragraph of narrative.
- One action per clip. Prompts asking for a sequence of events tend to produce neither.
- Name the look explicitly — “shot on 35mm, shallow depth of field”, “stop-motion”, “drone footage”.
- Iterate short. Generate at 4 seconds until the prompt is right, then re-run it longer.
What people use it for
A shot you need but have no footage for
An establishing shot, a texture, a mood insert. Faster than sourcing stock and specific to your brief.
Prompt low-angle shot of a red kite rising over wet sand at dusk, camera tilts up, 35mm
B-roll to cut between talking-head footage
Generate several 4-second variations and keep the two that cut cleanly. This is where the cheaper models earn their place.
Prompt hands close a laptop on a wooden desk, warm window light, shallow depth of field
Testing an ad idea before committing to a shoot
A rough generated version answers whether the concept reads at all, for a fraction of a day's production.
Prompt a runner crosses an empty bridge at sunrise, camera tracks alongside, breath visible
A clip that needs its own sound
Veo 3.1 generates synchronised audio in the same pass, which no other model here does.
Prompt rain on a tin roof at night, a single lamp swinging, distant thunder
Where it struggles
Text inside the frame comes out wrong more often than right. Prompts asking for a sequence of events tend to deliver neither half. Anything needing a specific real person, logo or place will be an approximation, not that thing.
Questions
Can AI make a video from just text?
Yes. Text to Video generates the footage from a written description alone, with no photo or video input.
How long can a text-to-video clip be?
1 to 16 seconds depending on the model. Veo 3.1 is fixed at 4, 6 or 8 seconds.
Can the video have sound?
Yes, with Veo 3.1, which generates synchronised audio. Any clip can also have a soundtrack added afterwards with Add Audio.
Which model is best for text to video?
Seedance 1.0 Pro for cost, Veo 3.1 for 4K and audio, Seedance 2.0 for complex scenes. You can switch per generation at no extra cost.
Does it work on Android?
No. The app is iPhone-only.