Wan text-to-video turns a scene description into a short generated clip. It works best when the prompt describes an observable action and camera plan, rather than a long collection of visual adjectives.
Generate a Wan video from text
Write six parts in this order: subject, action, camera, setting, light, exclusions.
A ceramic robot folds blue paper into a crane, one continuous slow clockwise camera orbit, clean wooden workbench, warm morning window light, physically plausible hands and paper, no captions, no logo.
If audio matters, name it plainly: soft paper folds, quiet mechanical ticking, no dialogue.
The model catalog is shared by the UI and the generation route. The server looks up credits, strips display-only tier suffixes, locks the matching duration and resolution, then charges atomically before provider submission. Unknown or canary-unverified models are rejected.
Last checked: August 12, 2026. WanVideoFree is an independent tool, not the official Wan site.
Use Wan 2.2 for the cheapest motion test. Use Wan 2.5 when native audio matters. Use Wan 2.7 for newer controls, longer provider durations, continuation, or editing workflows.
Long enough to specify subject, action, camera, environment, and constraints, but short enough that instructions do not compete. One clear paragraph is usually easier to debug than a page of adjectives.
After the server timeout boundary, a still-pending task is settled as failed and eligible credits are returned once. The browser does not own that refund decision.
Product event properties are sanitized to remove prompt, email, token, media, and URL fields. The inference provider still receives the prompt because it is required to perform the requested generation.
Two use-case walkthroughs build on this page: cinematic shots for depth-of-field and camera language, and short-form social video for 9:16 hook testing at the 8-credit tier.