In the first week of August 2026, China's two biggest tech companies shipped their flagship AI video models almost simultaneously. On July 31, ByteDance released Seedance 2.5; on August 6, Alibaba opened public beta for Wan 3.0. Both generate 30-second clips with native audio in a single pass — and both are being marketed as the model that finally moves AI video from "generate a clip" to "finish a project."
But read past the launch pages and it's clear they're built for different kinds of work. Wan 3.0 is the productivity model: it turns documents into videos, prices itself like a utility, and tops out at 1080p. Seedance 2.5 is the art-directed model: it swallows up to 50 reference materials, edits at timestamp level, and is already plugged into ByteDance's consumer apps. Which one you need depends on what you're making — and on one spec most comparison sites get wrong.
The spec sheet, for the record
| Spec | Wan 3.0 (Alibaba) | Seedance 2.5 (ByteDance) |
|---|---|---|
| Release | Aug 6, 2026 (public beta) | Jul 31, 2026; API Aug 7 |
| Max duration | 30s (2–30s per run) | 30s (4–30s), multi-round extension; 180s beta on web |
| API resolution | 480p / 720p / 1080p | 480p / 720p (web markets 4K upscale, not API) |
| Native audio | Yes — dialogue, SFX, music; can be disabled | Yes — 10+ languages, lip-sync |
| Reference inputs | 10 images + 5 videos + 5 audio (20 total) | 30 images + 10 videos + 10 audio (50 total) |
| Document-to-video | Yes: doc/xls/ppt/pdf/md + web links | Not available |
| Video editing | Instruction + reference editing | Timestamp-level local editing, green screen, clay render, camera edit |
| Extension | Built-in | Multi-round, high-fidelity |
| Prompt limit | 20,000 characters | No hard character limit |
| Billing | Per second + resolution | Per token (input-dependent) |
| Price (Beijing) | ¥0.3 / 0.6 / 1.2 per s (480/720/1080) | Token-based; estimated ~¥0.67/s 480p, ~¥1.51/s 720p |
| Open weights | No | No |
| Platforms | Alibaba first-party (preview), third parties rolling out | Jimeng, Doubao, Coze, CapCut, BytePlus ModelArk, global |
The one spec everyone gets wrong
Search "Wan 3.0 vs Seedance 2.5" and half the results will tell you both models output "native 4K." Both claims are wrong.
Wan 3.0's API has exactly three tiers — 480p, 720p, 1080p. There is no 4K. Seedance 2.5's situation is sneakier: the Dreamina web product advertises 4K and a 180-second beta mode, but the BytePlus API documentation lists only 480p and 720p, with a 30-second cap, and explicitly does not support 1080p or 4K. The fancy numbers live in the consumer apps; the API — the thing you'd actually build on — is 720p at most. So on pure API resolution, Wan 3.0 wins: it's the only one of the two that offers 1080p.
What the tests actually show
Atlas Cloud ran ten side-by-side tests with identical prompts on Wan 3.0, Seedance 2.5 and MiniMax H3. The honest conclusion: no single model won everything, and the differences were consistent and useful.
- In-frame text and clean product surfaces: Seedance 2.5. In test after test, Seedance kept on-screen text readable and product shots clean — the kind of thing that matters for ads and e-commerce.
- Long, over-specified multi-shot prompts: Wan 3.0. Wan handled the most complex storyboards in the batch — 1,700+ character prompts with ten timestamped shots — without losing the thread. It's built for detailed shot lists.
- Character identity and natural motion: neither was the clear winner. (MiniMax H3 took that category.) Both models drifted on some tests, especially on props across shots.
Independent Chinese testing (Photon Planet) found Wan 3.0's Chinese text rendering surprisingly stable — zero garbled characters in a PPT-to-video test — while Alibaba itself admits audio quality and text accuracy still need work overall. Seedance 2.5's own marketing emphasizes 10+ language support and reduced "uncontrolled subtitles and background music." Take both with a grain of salt until you run your own material.
Where Wan 3.0 shines
- Document-to-video is unique. Drop in a Word/Excel/PPT/PDF/Markdown file (≤100 MB / 50 pages) or a web link and get a narrated video. No competitor has this input channel yet — it's the single most practical feature for office and education work.
- 1080p API output. The only one of the two with a 1080p API tier.
- Transparent, cheaper pricing. Per-second billing you can quote before generating. 720p is ¥0.6/s — roughly a third to a half of Seedance's estimated per-second cost.
- Strong Chinese-language output and office-oriented workflows — decks, reports, product demos.
- Unified model covering text/image/reference-to-video, first/last-frame control, instruction editing and extension.
Where Wan 3.0 falls short
- Still preview / invitation-gated. The API is documented (Beijing and Singapore endpoints) but marked "preview" and "invitation only" — not every account has access yet. Third-party platforms like Atlas Cloud started onboarding on August 24, but it's not the free-for-all Seedance already has.
- Only 20 reference slots (10 images + 5 videos + 5 audio) vs Seedance's 50.
- Editing is less precise. Wan 3.0 has instruction and reference editing, but not Seedance's timestamp-level local editing or green-screen workflow.
- No 4K, and Alibaba concedes audio quality and text accuracy are still maturing.
- Prompt skill ceiling: expects timestamped shot lists (up to 20,000 characters) — more powerful, but a real learning curve.
Where Seedance 2.5 shines
- 50 multimodal references — 30 images, 10 video clips, 10 audio clips in one pass. For reference-heavy projects (multiple characters, scene kits, style packs) it's the most flexible input space of any video model today.
- Professional editing kit. Timestamp-level local edits (swap a background, fix a product, change an action), green-screen editing, clay-render (white-model) reference for locking composition, camera-perspective editing.
- Multi-round extension plus a 180-second ultra-long beta on the web — the longest single-generation experiments available.
- 10+ languages with native audio and lip-sync, aimed squarely at global and export-market content.
- Distribution is done. Live on Jimeng, Doubao Pro, Coze, Xiaoyunque, CapCut, plus BytePlus ModelArk API — with enterprise partners (XCMG, XPeng) already building on it.
Where Seedance 2.5 falls short
- API resolution caps at 720p. The 4K in the consumer marketing is an upscale on the web product; you cannot get 1080p or 4K through the API today.
- Token billing is a budget guessing game. No per-second price exists; your bill depends on input types, reference volume, retries and failed jobs. Analysts estimate ~¥0.67/s at 480p and ~¥1.51/s at 720p — comfortably more expensive than Wan 3.0 at the same tier.
- No document-to-video.
- Not open source, like Wan — no local deployment, no fine-tuning.
- Consumer-first posture: the flagship features (180s beta, 4K) live in the apps, not in the API you'd integrate.
The money question
These two models don't bill the same way, so don't compare a single number. Wan 3.0 charges per second and resolution: a 30-second 1080p clip is about ¥36; 720p about ¥18. Seedance 2.5 charges per token — with video input $0.0064/1K tokens, without video input $0.0107/1K tokens — and the final cost depends on your prompt, references and retries. Third-party estimates put a comparable 30-second 720p clip around ¥45, roughly 2.5× Wan 3.0's price at the same tier. Seedance may still win per approved clip if its higher hit rate saves you rerolls — that's a question only your own test runs can answer.
So, which one?
Pick Wan 3.0 if your work is document-driven (decks, reports, lessons), you need 1080p through an API, you want a budget you can quote before you press generate, or you mostly produce Chinese-language content.
Pick Seedance 2.5 if your projects are reference-heavy (many characters, props, style packs), you need serious editing (local fixes, green screen, white-model pre-vis), you're shipping multi-language content, or you already live in the ByteDance ecosystem.
And honestly, the two are complementary. Wan 3.0 is the "get it done" model for office and education production; Seedance 2.5 is the "art-direct it" model for film-like, multi-reference work. Before committing, run the same prompt and references through both — the spec sheet answers the easy questions, but your footage answers the real ones.
Try Wan 3.0 and Seedance 2.5 side by side on wan3video.io — same prompt, same references, and you can see the difference in minutes.
