Text-to-Video with Alibaba Wan 3.0
Alibaba's Wan 3.0 generates up to 30 seconds of video in a single take, with native synchronized audio — dialogue, sound effects, and music — produced in the same pass as the picture. Pick any duration from 2 to 30 seconds, output at 480p, 720p, or 1080p, and choose from five aspect ratios or let the model frame the shot automatically.
What makes Wan 3.0 different
Single takes up to 30 seconds
Most video models top out at 10 or 12 seconds per generation and force you to stitch clips together. Wan 3.0 generates any length from 2 to 30 seconds in one pass — long enough for a complete ad spot, a full scene, or an entire product walkthrough with no cuts, no extends, and no continuity drift between segments.
Native audio — and it's the cheaper setting
Wan 3.0 generates dialogue, sound effects, and music synchronized to the picture in the same pass, steered by your prompt. Unusually, leaving audio on is also the cheaper configuration: Scenetra routes every render to the cheapest provider eligible for your exact settings, and audio-on renders qualify for a lower rate — $0.025 per second at 480p versus $0.04 with audio switched off. Unless you specifically need a silent clip, keep the default.
Three resolutions, one price ladder
480p at $0.025 per second is the drafting rung — a full 30-second take costs $0.75. The 720p default runs $0.05 per second, and 1080p finals cost $0.10 (audio-on rates). A practical workflow on a Scenetra board: iterate long takes at 480p until the pacing and audio land, then flip the resolution select to 1080p for the final render — same node, same wiring.
Billed at cost in a visual workflow
Wan 3.0 runs as a node in Scenetra's node-based editor, so you can chain it with image models — generate a style frame, refine it, then feed the look into your video prompts on one canvas. Every generation is billed at provider cost with 0% markup, pay-per-generation with no subscription gate, and new accounts get free welcome credits plus a 7-day trial.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Try Wan 3.0 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the video: subject, motion, camera work, and audio direction. Native synchronized audio is generated in the same pass. | — |
| Resolution | select | Output resolution. 480p is the budget rung; 1080p is the premium rung.480p · 720p · 1080p | 720p |
| Aspect Ratio | select | Aspect ratio of the generated video. Auto lets the model choose the best framing for your prompt.Auto · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | Auto |
| Duration | select | Length of the generated video in seconds. Wan 3.0 generates any length from 2 to 30 seconds in a single take.2 · 3 · 4 · 5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 · 16 · 17 · 18 · 19 · 20 · 21 · 22 · 23 · 24 · 25 · 26 · 27 · 28 · 29 · 30 | 5 |
| Audio | boolean | Generate synchronized audio (dialogue, SFX, music) in the same pass. Leaving it on is also the cheaper setting — see pricing. | true |
Pricing
From $0.025 per second
Pay per generation — only pay for what you use.
480p (with audio)
$0.025/s
per second
480p (audio off)
$0.04/s
per second
720p (with audio)
$0.05/s
per second
720p (audio off)
$0.08/s
per second
1080p (with audio)
$0.10/s
per second
1080p (audio off)
$0.16/s
per second
What a video costs
| Duration | 480p (audio) | 480p (muted) | 720p (audio) | 720p (muted) | 1080p (audio) | 1080p (muted) |
|---|---|---|---|---|---|---|
| 5 seconds | $0.125 | $0.20 | $0.25 | $0.40 | $0.50 | $0.80 |
| 10 seconds | $0.25 | $0.40 | $0.50 | $0.80 | $1.00 | $1.60 |
Use cases
Marketing & ads
- Complete 30-second spots in one take
- Product teasers with synced sound
- Vertical 9:16 social ads
- Multi-length campaign variants
Social content
- Sound-on shorts and reels
- Long-form single-shot storytelling
- Ambient mood loops
- Square and portrait cuts from one prompt
Film & previz
- Full-scene previz without stitching
- Dialogue and SFX scratch tracks
- Establishing shots with ambience
- Cheap 480p blocking passes
Education & explainers
- Narrated explainer segments
- Step-by-step process clips
- Concept visualizations with sound
- Course intro sequences
Related models
Seedance 2.5
ByteDance's Seedance 2.5 generates cinematic videos up to 30 seconds from text prompts with synchronized audio — voice, sound effects, and background music — built in by default. The newest Seedance generation, with mp4 or high-fidelity mov output for multi-round editing pipelines.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →MiniMax H3
MiniMax H3 is the newest generation of the Hailuo video lineage: 5-15 second clips with native stereo audio — dialogue, music, and sound effects rendered in the same pass as the picture. Generate at 480P or 768P natively, or step up to 2K and 4K upscaled output, in six aspect ratios from cinematic 21:9 to vertical 9:16.
View model →Frequently asked questions
What is Wan 3.0?+
Wan 3.0 is Alibaba's video generation family. The text-to-video mode turns a prompt into a video of 2 to 30 seconds in a single take, with native synchronized audio — dialogue, sound effects, and music — generated in the same pass. It outputs at 480p, 720p, or 1080p in five aspect ratios plus an auto framing mode. On Scenetra it runs as a node in the visual workflow editor.
How much does Wan 3.0 cost?+
With audio on (the default), $0.025 per second at 480p, $0.05 at 720p, and $0.10 at 1080p — so a 10-second 720p clip with sound is about $0.50, and a full 30-second 1080p take is $3.00. With audio switched off the rates are higher: $0.04, $0.08, and $0.16 per second. Everything is billed at provider cost with 0% markup.
Why do muted renders cost more than renders with audio?+
Scenetra always routes your job to the cheapest provider that supports your exact settings. The audio-on configuration is eligible for a lower rate than the audio-off one, so switching audio off narrows the routing options and lands on a pricier path — $0.04 per second at 480p instead of $0.025. If you don't need the soundtrack, it's still cheaper to generate with audio and mute the file yourself.
Can Wan 3.0 really generate 30-second videos?+
Yes — duration is selectable at every second from 2 to 30, and the whole clip is generated as one continuous take rather than stitched segments. That makes it one of the few models that can produce a complete 30-second ad or scene in a single generation, with the audio track running the full length.
Does Wan 3.0 generate audio?+
Yes, natively. Wan 3.0 produces dialogue, sound effects, and music synchronized with the picture in the same generation pass, and you can direct the audio in your prompt alongside the visuals. Audio is on by default, adds nothing to the price — in fact the audio-on rate is lower — and can be disabled with a toggle if you want a silent clip.
Is Wan 3.0 free to try online?+
You can try it free through Scenetra: new accounts get free welcome credits and a 7-day trial that cover your first generations. After that it's pay-per-generation — about $0.125 for a 5-second 480p clip with audio — with no subscription gate and nothing to cancel.
What's the difference between Wan 3.0's text-to-video, image-to-video, and reference-to-video modes?+
Text-to-video generates entirely from a prompt. Image-to-video animates a still image you upload as the starting frame, at the same prices. Reference-to-video takes 1-10 reference images and keeps those people, objects, or styles consistent in the output; it's billed at a higher flat rate ($0.04 to $0.16 per second depending on resolution). All three share the 2-30 second single-take range and native audio.
Start creating with Wan 3.0
Use Wan 3.0 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free