Coming soon to A2E
Seedance 2.5: Professional Multimodal AI Video
ByteDance’s new-generation professional multimodal video creation model. Four upgrades in one release — long narrative, strong reference, precise editing and multilingual. Built for real production work, so footage becomes reusable, deliverable and repeatable at scale.
Four upgrades in Seedance 2.5
Every figure below is taken from ByteDance’s official release notes. Actual results on A2E will be confirmed once the model goes live here.
Long narrative
A single shot now runs 30 seconds instead of 15. That is long enough to carry a complete emotional arc without stitching clips together in post. High-fidelity temporal extension can also grow short footage into a longer cut while keeping character, scene and camera consistent.
Strong reference
Up to 50 reference assets in a single generation, against 12 in Seedance 2.0. Feed a cast, a set of locations, reference camera moves and brand music at once; the model reads them together and returns a cut that matches the brief. Faces, expressions and wardrobe stay stable across shots.
Precise editing
Feed an existing cut as @video1, describe the change in the prompt, and the model rewrites only that part while everything else holds. Swap a product, restyle the footage, adjust an actor’s age or expression, remove a watermark, or replace the background music.
Multilingual
More than ten languages natively, so creators can describe intent in their own language without a translation round trip. One shoot can also be released as several localised voice-over versions — the clip on the left is the same scene delivered in English, French, Japanese, Spanish, Korean and Hindi.
What changes from Seedance 2.0
ByteDance is explicit that 2.5 is not a generational jump the way 2.0 was over 1.5. Version 2.0 already delivered the core capability breakthrough; 2.5 is a systematic reinforcement aimed at real production work.
| Capability | Seedance 2.0 | Seedance 2.5 |
|---|---|---|
| Single-shot duration | 15s | 30s |
| Reference assets per generation | 12 | 50 |
| — Images | 9 | 30 (up to 4K) |
| — Video clips | 3 | 10 |
| — Audio clips | 3 | 10 |
| Total video / audio reference length | 15s | 30s |
| Audio-only reference | Not supported | Supported |
| Native languages | — | 10+ |
Know which task you are running
ByteDance gives this its own section for a reason: reference, editing and extension get mixed up in practice, which misplaces the prompt, cites the wrong assets and costs you control and stability. Pick the task first, then write the prompt for it.
Reference generation
When: you have no existing cut. Build a new video from zero using images, video or audio as reference.
Prompt shape: what you want generated, in plain text, plus @video N @image N @audio N to cite each asset.
Locked parameters: none. You choose aspect ratio and duration yourself.
Video editing
When: you already have a cut and want part of it rewritten, replaced, added to or removed.
Prompt shape: the change you want, plus @video1 for the source cut, optionally @image N or @audio N.
Locked parameters: aspect ratio follows the source exactly, so nothing is stretched or cropped. Output length tracks the source within roughly 0.3 seconds. MOV is the recommended output format. Keep the input under 20 seconds.
Video extension
When: you already have a cut and want it continued in the same look, pacing and style.
Directions: forwards from the last frame, backwards from the first frame, or filling the gap between two or three clips.
Locked parameters: aspect ratio follows the source. The length of the new segment is yours to set — state it in the prompt. MOV is recommended here too.
First frame and first-last frame
Give one image as the opening frame, or two images as opening and closing frames. Aspect ratio follows the first frame; if the closing frame has a different shape it will be stretched to match, so supply frames of equal proportion. Duration stays under your control.
Fifty references, and how many to actually use
The ceiling is 50 assets — 30 images, 10 video clips and 10 audio clips. ByteDance is equally clear that hitting the ceiling is not the goal. The ratios below come from their own large-scale evaluation.
Hard limits
- Images: up to 30, 4K or smaller, 30 MB each. JPEG, PNG, WebP, BMP, TIFF, GIF, HEIC, HEIF.
- Video: up to 10 clips, 480p to 4K, 200 MB each. MP4 or MOV. Each clip 2–30s, 30s total.
- Audio: up to 10 clips, 15 MB each. MP3 or WAV. Each clip 2–30s, 30s total.
- Video and audio budgets are counted separately, not combined.
What actually works
- Subject carried by video or audio: 1–5 assets works well. 6–10 is worth trying but stability drops and you may need several attempts.
- Each such clip is best at 5–10 seconds. Two seconds carries too little to identify; thirty seconds dilutes the features and adds noise.
- Subject carried by stills: 1–8 images works well, 9–12 is worth trying.
- Up to 5 subjects, single or multi-view images are both fine. Beyond 5, single-view is steadier — split the angles into separate images rather than packing them into one.
A “subject” is any core element you want held stable through the whole cut — a character, a hero product or prop, the location, the overall style. In “someone in a forest shooting a rabbit with a bow”, the person, the bow and the rabbit are three subjects.
Writing prompts for 2.5
Prompt engineering changed in this release and the model now expects more from the text. ByteDance’s guidance is to think like a director and write a structured prompt in four parts.
- Cite your assets. Number every image, video and audio file in upload order and say what each one is for — who is the likeness, which is the voice, which supplies the motion, which sets the location.
- One-line summary. Subject, place, event, genre or style, and any special camera work.
- Beat-by-beat detail. Walk through the shots in order — timestamps or “shot N” both work. Describe framing, camera move, action, dialogue and sound effects. Write in the positive.
- Close it out. Add the details that run through the whole piece: lens and camera position, environment, voice, atmosphere.
The official prompt-tuning skill
ByteDance strongly recommends their own prompt-engineering skill for Seedance 2.5. It installs locally through npx and runs inside your own AI chat window — no third-party service involved.
npx --yes skills@latest add \
"https://arkdocs.tos-cn-beijing.volces.com/skills/" \
--skill sd25-pe \
--yes
Then type /sd25-pe followed by your draft prompt.
Negative control is supported
You can rule things out as well as ask for them. Subtitles and audio have their own switches — “no on-screen subtitles”, “no BGM, ambient and action sound only” — and audio can be steered separately as sound effects, background music or dialogue.
A few things that quietly go wrong
- Timestamps read best in one-second units. Naming a frequency inside a window — “shake the head three times in one second” — is not recommended.
- Storyboard grids work up to about 15 panels. Stick figures or line art beat polished art, and heavy on-panel text hurts.
- Grey-box previz works better coarse than fine — simple geometry assembled roughly reads more clearly than a detailed model.
- Map subjects to assets explicitly. Writing a name only on the image and then using that name in the prompt invites the model to mix people up.
What you can set per generation
These are the controls ByteDance exposes for the model today. Once Seedance 2.5 is live on A2E, whatever A2E offers is what applies.
- Mode — reference generation, first / first-last frame, video editing, or video extension. Pick this first: it decides which of the settings below stop being yours to choose.
- Resolution — 480p or 720p.
- Aspect ratio — 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, or let the model decide from your assets.
- Duration — set it in seconds, or hand it to the model.
- Output format — MOV is recommended for editing and extension jobs, where it holds colour, brightness and audio-video sync better.
- Audio — on or off.
- Batch size — up to 8 clips in one go.
Docs and API
Building Seedance into your own product? Endpoints, parameters, authentication and billing are all in the A2E developer docs, and MCP access is available there too.
Why run Seedance on A2E
Start free, no subscription
Generate without a monthly plan. Seedance 2.0 is already live, so you can build the workflow now and move to 2.5 when it lands.
One place for every model
Seedance, Kling, Wan, Sora, Veo, Seedream and more sit behind one account and one balance. Compare them on the same brief instead of paying for several tools.
Built for delivery
Digital humans, lip sync, voice cloning, face and head swap and upscaling live alongside generation, so a clip can go from prompt to finished asset in one place.
Related Seedance pages on A2E
- Seedance 2.0 video generator — the version live on A2E today.
- Seedance 2.0 Mini — the faster, lower-cost option when volume matters more than polish.
- Seedance 2 prompts — prompt patterns that carry over to 2.5.
- Image to video and text to video — the two entry points most people start from.
- Compare against alternatives: vs Kling, vs Veo, vs Sora.