Now live · BETA
Wan 3.0: All-in-One AI Video
Tongyi Wanxiang’s new-generation all-in-one video model, live on A2E today. A native 30-second single take, four generation modes in one place, up to 20 reference assets, and documents or web pages read straight into the video.
What the model puts out
Every clip below comes from Tongyi Wanxiang’s own launch demos, so you can judge the picture quality before spending anything.
Five things it does well
The five headlines and their numbers are Tongyi Wanxiang’s own; the clips are their demos. What you get on A2E is whatever your own generation returns.
Omni-creation — up to 20 assets at once
Generate with up to 20 reference assets in a single job, including complex document and web page parsing — doc, xls, ppt, pdf and md all read as reference input. A product spec or a competitor’s page can become the source material for a video without being rewritten first.
Native 30 seconds, not stitched
A single take now runs 30 seconds natively, which allows more complete storytelling and greater creative freedom, with an intelligent duration setting that picks a length from your prompt. One clip can hold a full emotional arc, so continuous camera moves and one-take shots have room to play out.
Pixel-perfect consistency, built for delivery
It reproduces reference details precisely, which is aimed at production work rather than experiments: characters, props, voices, spatial relationships and overall style hold steady from shot to shot. Whether the same face survives a cut is the line between a demo and something you can hand over.
Picture and sound in one pass
Realism, visual texture and sound design all move up together for a stronger audiovisual impact. Audio generation is a switch on the generation form, so picture and sound come out together — no separate scoring pass, and no syncing afterwards.
Precision editing — change one thing, keep the rest
Video editing is more precise in this release, with support for instruction-based and reference-based editing. Say which part should change — the picture, the action, a line of dialogue — and the rest of the cut stays as it was instead of being regenerated from scratch.
Pick the mode before the settings
A2E opens four modes for Wan 3.0. Which one you choose decides what you need to prepare and which settings your assets take out of your hands.
Text to Video
Prompt only, no upload. Describe the scene, what the subject does, how the camera moves and the mood you want, and the model builds it from nothing. The fastest way to try an idea out.
Image to Video
Upload one image as the first frame and the model generates what follows from your prompt. It also accepts an audio upload, so the result comes back with sound. If you already have a key visual or a product shot, this is the shortest route.
First & Last Frame to Video
Upload an opening frame and a closing frame separately and the model fills in the transition between them; a swap button flips the two. This is the steady choice for expression changes, pose changes and scene transitions — anything where you need to control exactly where the shot starts and ends.
Reference to Video
Upload several reference images, videos or audio clips, or point at a document or web page, and Wan 3.0 reads them together to steer the result. Keep the reference video under 15 seconds in total: the system adjusts the output so input and output together stay within 30 seconds.
Three steps to your first clip
This is the actual order of operations on A2E. Nothing extra to activate first.
1. Choose the mode
Decide first whether a prompt alone, one image, a first and last frame, or multi-modal reference material will drive the video. The four modes are described above.
2. Upload and write the prompt
Add whatever that mode needs — images, video, audio, or a document or web page — then describe the action, the camera move and the atmosphere you are after. The prompt field takes up to 5,000 characters, and specific beats concrete.
3. Set the parameters and generate
Pick duration and resolution, then hit generate. The result lands in My Result, where you can carry it on into upscaling, subtitle removal and the rest of the toolbox.
What you can set per generation
These are the controls Wan 3.0 exposes on A2E today — enough to judge whether it clears your delivery bar.
| Setting | What you get |
|---|---|
| Duration | 3s, 5s, 10s, 15s or 30s. Default 5s. In Reference to Video the available range moves with the length of the reference clip you uploaded. |
| Resolution | 480P, 720P or 1080P. Higher costs more credits — rough the idea out at 480P, finish at 1080P is the cheap way to work. |
| Input images | JPG or PNG, 20 MB or less each, with both width and height between 240 and 7,680 pixels. |
| Prompt | Up to 5,000 characters. A negative prompt rules things out (“blurry, low quality, deformed”), and prompt expansion lets the AI fill your description out for you. |
| Generate audio | On by default, and you can turn it off. |
| Number of outputs | Several clips from one submission. |
| Random seed | Ultra plan only. Reproduces the same result from the same inputs. |
Four things that save a retry
Start from good material
Sharp, well-composed images and video give better results; blurry or low-resolution input does not. The quality of what goes in is the ceiling on what comes out, and that matters more than any setting on this page.
Name the action
Concrete movements in the prompt — “walks slowly”, “nods gently” — land far more accurately than adjectives. Vague mood words give the model nothing to hold on to.
Use first-and-last frame on purpose
That mode is built for expression changes, pose changes and scene transitions — anywhere the start and end of the shot have to be exact. When you already know where the shot goes, do not leave it to a text prompt.
Budget the reference length
In Reference to Video the reference clips must total 15 seconds or less, and the system trims the output so input plus output stays within 30 seconds. Longer references leave you a shorter film — do that arithmetic before you upload.
Docs and API
Building Wan 3.0 into your own product? Endpoints, parameters, authentication and billing are in the A2E developer docs, and MCP access is available there too.
Why run Wan 3.0 on A2E
Start free, no subscription
Generate without a monthly plan. Wan 3.0 runs on your existing credits, and there is nothing extra to activate before your first clip.
One place for every model
Wan, Seedance, Kling, Sora, Veo, Seedream and more sit behind one account and one balance. Run the same brief through several of them instead of paying for several tools.
Built for delivery
Lip sync, voice cloning, face and head swap, subtitle removal and upscaling live next to generation, so a clip can go from prompt to finished asset without leaving the site.
Related AI video models on A2E
- Seedance 2.5 — ByteDance’s professional multimodal model, also live on A2E. Compare the two on the same brief.
- Wan 2.7 and Wan 2.6 — the previous Wan generations, still fast and economical options.
- Kling 3.0 and Sora 2 — alternative video models worth running the same prompt through.
- Lip sync and video to audio — finishing tools for generated footage.