Domo AI · Text to video

Domo AI Text to Video

Write a scene and get an animated clip. No footage needed

Describe a subject, an action, a place, a style and a camera move, and DomoAI generates a video from your words. Use DomoAI 2.4 for quick anime and stylized clips with Relax mode, or Seedance 2.0 for clips up to 15 seconds in up to 1080p with native audio.

4–15 sper clip
Up to 1080poutput
Any languagefor prompts
Prompt

A girl with a red umbrella walks through a rainy neon street at night, anime style, slow tracking shot, reflections on wet pavement.

Video scene generated from a text prompt in Domo AI
How it works

From text to video in 4 steps

Beginner’s guide →
  1. 01

    Write the prompt

    Subject, action, environment, style and camera. Any language works.

  2. 02

    Pick a style

    Japanese anime, realistic, pixel art, 90s aesthetic, or the default.

  3. 03

    Choose settings

    Model, duration and aspect ratio (1:1, 16:9, 9:16 and more).

  4. 04

    Generate & review

    About 2–3 min for 5 s and 3–5 min for 10 s on DomoAI 2.4.

Models

Which text-to-video model to use

SpecDomoAI 2.4 FastDomoAI 2.4 AdvancedSeedance 2.0 FastSeedance 2.0
Length5 or 10 s5 or 10 s 4–15 s 4–15 s
ResolutionNot specifiedNot specified480P–1080P480P–1080P
Native audio—— On/off On/off
Relax mode Standard+ Standard+ Fast only Fast only
Aspect ratios1:1, 16:9, 9:16…1:1, 16:9, 9:16…3:4, 4:3, 9:16, 16:9, 1:13:4, 4:3, 9:16, 16:9, 1:1
Best forQuick drafts & iterationsPrecision & consistencyFast longer clips with soundLonger, higher-res scenes with sound

Neither model family outputs 4K directly, so use the Video Upscaler. For multi-reference scenes with Seedance 2.5, use Omni Reference.

Prompt structure

The 5-part video prompt

DomoAI recommends this order: Subject + Action + Environment + Style + Camera movement. Write it as clear, complete sentences.

  1. Subjectwho or what
  2. Actionone clear movement
  3. Environmentwhere and when
  4. Styleanime, realistic, pixel…
  5. Camerapush-in, orbit, tracking
Example, broken down

A boy playing keyboard on a hill at sunset, soft anime style, pastel colors, slow camera push-in.

1 Subject2 Action3 Environment4 Style5 Camera

Video adds two things an image prompt doesn’t need: action and camera movement. Keep both to one each per clip.

Interactive

Text-to-video prompt builder

Fill in the scene, choose what matters most, and get a ready prompt with a model suggestion. Switch to multi-shot to write a CAM sequence.

Style
Camera
Duration
Aspect
What matters most?
Suggested model
Your prompt

        
Advanced prompts

Techniques from DomoAI’s guide

Prompt writing guide →

CAM shot labels

Label shots with time ranges to direct a sequence inside one clip.

CAM 1 (0–3 s): wide shot of the city. CAM 2 (3–5 s): close-up of her face.

Film language

Use cinematography terms for a cut-series look: lens, depth of field, color grade.

35mm lens, shallow depth of field, teal-and-orange grade

AI Optimize

Turn it on to expand a short idea into a detailed prompt automatically, then edit the result.

JSON structure

For complex scenes, a structured prompt keeps each element separate.

{"subject":"fox in a scarf",
 "action":"leaps over a log",
 "scene":"snowy forest, dawn",
 "style":"anime",
 "camera":"slow tracking shot"}
Styles

Built-in looks

Prompts by style →
Japanese anime style video

Japanese anime

Clean line art and cel shading for stories and fan content.

Realistic cinematic video

Realistic

Cinematic, photo-like scenes for ads and B-roll.

Pixel art style video

Pixel art

Retro game looks for gaming channels and intros.

90s aesthetic video

90s aesthetic

Nostalgic color, grain and VHS vibes.

Which tool?

Text to video vs image or video input

Text to video is the most open-ended option. When you need a specific look or movement, start from an image or a clip instead.

Start fromControlBest for
TextLowest: the AI invents the lookIdeas, storyboards, B-roll
ImageThe first frame is fixedExact characters and compositions
VideoMotion is fixedRestyling real footage
ReferencesHighest: you assign rolesConsistent characters across shots

Tips for better clips

  • One subject, one action and one camera move per clip
  • Test at 5 s before you spend credits on longer clips
  • Draft on DomoAI 2.4 in Relax mode, then do the final on Seedance
  • Generate a character image first and use Image to Video for consistency
  • Use the same style line across clips in a series

Limitations

  • Short clips: 5–10 s on DomoAI 2.4, up to 15 s on Seedance 2.0
  • Less predictable than image or video input
  • Characters can change between separate generations
  • Fast, complex actions and hands can warp
  • No direct 4K: upscale afterwards
All limitations →
For developers

Text to video in the Domo API

Call POST /v1/video/text2video with a prompt, 1–10 seconds and a model. Faster models cost 2 credits per second and advanced models 5, at $0.02 per credit.

API guide
API modelCredits / s5 s10 s
t2v-2.4-faster2$0.20$0.40
t2v-2.4-advanced5$0.50$1.00
What is Domo AI Text to Video?

DomoAI’s text-to-video tool turns a written description into a video clip with no footage or images needed. You describe the subject, action, setting, style and camera, choose a model and length, and DomoAI generates the clip.

Which models does Text to Video use?

DomoAI 2.4 Fast and 2.4 Advanced (5 or 10 seconds, with Relax mode on Standard and above) and Seedance 2.0 and 2.0 Fast (4–15 seconds, 480P–1080P, optional native audio). Seedance 2.5 is available through Omni Reference.

How long can a text-to-video clip be?

5 or 10 seconds on DomoAI 2.4, and 4–15 seconds on Seedance 2.0. Build longer videos by joining several clips.

How long does generation take?

On DomoAI 2.4, about 2–3 minutes for a 5-second clip and 3–5 minutes for 10 seconds. Busy times and longer, higher-resolution clips take longer.

How should I write a text-to-video prompt?

Use Subject + Action + Environment + Style + Camera movement, in clear complete sentences. Keep to one action and one camera move per clip. Prompts work in any language.

Can I make videos with sound?

Yes. Seedance 2.0 can generate native audio, which you can switch on or off. You can also add voiceovers with Talking Avatar or in your editor.

Is Text to Video free?

New accounts get 30 starter credits to test it, with daily limits per tool on the free plan. Paid plans add more credits, and Standard and above include Relax mode on supported models such as DomoAI 2.4.

How much does text to video cost through the API?

With t2v-2.4-faster it costs 2 credits per second ($0.20 for 5 s), and with t2v-2.4-advanced 5 credits per second ($0.50 for 5 s), at $0.02 per credit.

Write it. Watch it move.

Turn your first scene into a video with 30 free starter credits.

Start Creating Free Build a prompt