Write a sentence. Get a video with sound. No footage, no stock library, no editing software.
Text-to-video builds the entire scene from your description — which makes it the right tool when you have an idea but no image, and the wrong tool when you need a specific thing to appear. If you already have the subject as a picture, image-to-video will hold on to it far more faithfully.
Being honest saves you credits:
A usable prompt usually names four things: subject, action, setting, camera.
A weathered fishing boat rocking gently, moored in a harbour at sunrise, seagulls overhead, slow push in from the water
Compare that with boat in harbour — same idea, far less to work with.
Then:
Starter credits on signup, free generations on a lighter model, Wan 2.5 at full quality on paid tiers.
The reason for that split: Wan 2.5 costs about $0.05 per second at 480p and $0.15 at 1080p in real GPU time. A site offering unlimited free 1080p video is burning money it doesn't have. We'd rather tell you where the line is than surprise you with it later.
Yes, within starter credits. Free generations use a lighter model; Wan 2.5 full quality is paid. No card needed to try.
Roughly 1–3 minutes for a 5–10 second clip, depending on resolution and queue.
Yes — Wan 2.5 generates native audio alongside the picture, already synced. Most models leave you with a silent clip.
16:9 for YouTube and landscape, 9:16 for Shorts/Reels/TikTok, 1:1 for feed posts.
If you can produce or find the key image, use image-to-video. It preserves your subject and gives dramatically more control. Use text-to-video when the scene doesn't exist yet.
Yes on paid plans. Standard restrictions apply: nothing illegal, no impersonation of real people, no content designed to deceive.