Skip to content

Video generation -

Reference, synced 2026-06-13.

flowchart TD
  n0["Models"]
  n1["Overview"]
  n2["Products"]
  n3["Solutions"]
  n4["Pricing"]
  n5["Resources"]
  n6["Partners"]
  n7["Support"]
  n8["Language"]
  n0 --> n1
  n1 --> n2
  n2 --> n3
  n3 --> n4
  n4 --> n5
  n5 --> n6
  n6 --> n7
  n7 --> n8

Free access Accelerate Delivery with Fixed-Cost Agentic CodingWatch how it works

Models

Empowering AI innovation for both enterprises and developers with Alibaba Cloud’s best-in-class Qwen models, AI-native apps, and AI solutions.

Alibaba Cloud Model Studio \ Enterprise-grade large model service and application development platform.

Try Visual Model \ Supports image understanding, image generation, and video generation.

Models

HappyHorse-1.0-T2V \ Cinematic creative generation, ultimate dynamic details Qwen3-VL-Plus \ Native VL, spatial reasoning, 1M-context video analysis Wan2.7-VideoEdit \ Supports both localized and global editing with prompt

Qwen3.6-Plus \ Native multimodal, 1M context, agentic coding Wan2.7-Image-Pro \ Interactive editing, long-text rendering, precise prompt following Qwen-Plus \ Balanced intelligence, efficient inference, production-ready performance

Qwen-Image-2.0 \ Professional infographics, exquisite photorealism Z-Image-Turbo \ Ultra-fast image generation, high throughput, cost-optimized inference Qwen3-Coder-Next \ Multi-turn tool interactions, future-ready development support

Wan2.7-T2V \ High-fidelity T2V, 15s duration, advanced camera control Wan2.7-I2V \ Cinematic I2V with emotional depth and visceral impact Wan2.7-R2V \ Up to 5 mixed image/video inputs and audio timbre cloning

GenAI Application

Qoder \ Intelligent coding assistant, available for enterprise-dedicated deployment. Qoder CN \ AI-powered coding assistant that boosts developer productivity with intelligent code completion, AI chat, multi-file editing, and task automation.

AI Service

Model Experience \ Experience full-scale, multimodal model capabilities online. Platform for AI \ An AI-native algorithm engineering platform for end-to-end modeling, training, and inference service deployment. Fine-tune Video Generation Model \ Customize Wan’s text-to-video capabilities through model fine-tuning to meet your unique requirements.

AI Use Case

AI Savings Plan Hot \ Save up to 47% on AI costs. Limited-time offer tailored to your usage. AI Video Creation \ Elevate your professional video production with Wan 2.6.

AI Token Plan \ One plan. Multiple models. Big Savings with a Fixed Subscription. AI Image Creation \ All-in-one creative suite for copywriting, image generation, and poster design.

Overview

As a global full-stack AI leader, Alibaba Cloud aims to make computing accessible to everyone and help worldwide customers accelerate innovation.

Why Alibaba Cloud

About Alibaba Cloud \ AI Powered Cloud Technology Our Global Network \ Explore our global presence and deployment regions around the world Our Global Offices \ With offices in 4 continents, we're always close to where it matters.

Asia Accelerator \ Accelerate Success in Asia with Alibaba Cloud Go Global \ Benefits of our Global Alliance Trust Center \ Empowering enterprises with a secure, compliant, and globally trusted cloud infrastructure

Customers and Insights

Olympic Games \ Alibaba Cloud Powers Olympic Games with AI-powered cloud technology Case Studies \ Learn how customers are scaling their businesses on Alibaba Cloud Analyst Reports \ Learn what the top industry analyst firms are saying about Alibaba Cloud

What's New

Events and Webinars \ Quick access to upcoming and on-demand events Product Updates \ Stay informed of the latest innovations Press Room \ Latest news and media releases

Products

Featured ProductsAI & Machine Learning Computing Container Storage Networking & CDN Security Middleware Database Analytics ComputingMedia ServicesEnterprise Services & Cloud CommunicationDomain Names and WebsitesEnd User ComputingServerlessDeveloper ToolsMigration & O&M ManagementApsara Stack

Alibaba Cloud Model Studio \ Supercharge your AI journey effortlessly with industry-leading GenAI models ApsaraDB RDS \ Store and manage your business data, with automated monitoring and backups Certificate Management Service (Original SSL Certificate) \ Create a safe and secure connection between your website and users

Elastic Compute Service (ECS) \ Host websites anywhere and scale enterprise workloads Container Service for Kubernetes (ACK) \ Run and scale containerized applications on managed Kubernetes infrastructure Object Storage Service (OSS) \ Store large amounts of data in the cloud and access it anywhere, anytime

Simple Application Server (SAS) \ All-in-one services for fast deployment Elastic IP Address (EIP) \ Manage your public IPs independently to improve internet network quality Domain Names and Website \ Get the perfect domain name to suit your every need

Solutions

Solutions by Industry Technical Solutions AI WebsitesNetworking Security and ComplianceData and AnalyticsEnterprise Service and ApplicationCloud MigrationCloud NativeHybrid CloudSMB solutions

Financial Services \ Innovate faster with Alibaba Cloud Games \ Grow your game rapidly with high global availability

New Retail \ Alibaba Cloud enables digital retail transformation to fuel growth and realize an omnichannel customer experience throughout the consumer journey. Media and Entertainment \ Ready your content for today's media market with a digitalized media journey

Supply Chain \ Power your supply chain with intelligent, efficient, and reliable solutions Sports \ Digitizing the sports industry with intelligent tech

Sustainability \ Achieve a sustainable future with low-carbon and energy-efficient technologies

Pricing

Flexible options like pay-as-you-go and clear billing rules to meet diverse business needs.

Overview & Tools

Pricing Calculator \ Get an instant pricing estimate based on your usage and needs Free Trial \ Try our 80+ cloud products for free.

Pricing Options \ Get the most out of Alibaba Cloud with flexible pricing

Optimize your cost

Migrate & Save \ Superior Performance At Lower Pricing. Save up to 50%. Promotion Center \ Unlock the latest Alibaba Cloud offers & promos

Resources

Official documentation, extensive tools, training resources, and a community to grow and innovate in the cloud.

Technical Resources

Documentation \ Product guides and FAQs Architecture Center \ Design reliable, secure, and efficient cloud architecture. Intelligent Solution Explorer \ Find the right solution for you, powered by AI

Blog \ Latest cloud insights and developer trends Whitepapers \ Research that explores the how and why behind our technology

Training&Certification

Alibaba Cloud Academy \ Build cloud skills and earn certifications with expert-led training.

Developer Hub

Alibaba Cloud Project Hub \ Explore real-world projects built by developers using our platform. Our Developer MVPs \ Celebrating the developers who lead, build, and inspire our community

Partners

Partner-first strategy offering collaborative product, sales, and service models, plus high-quality partner solutions that complement Alibaba Cloud’s capabilities.

Marketplace

AI Alliance for ISVs \ Partner with us to build and grow AI solutions together ISV Benefits \ Unlock resources, market access, and go-to-market support as an ISV partner

Alibaba Cloud Marketplace \ Explore ready-to-deploy solutions from our partners and ISVs

Find a Partner

Partner Hub \ Find your ideal partner in no time

Become a Partner

Partner Network \ A partner portal for Alibaba Cloud Channel, Technology, MSP partner and other partner programs

Support

Full-lifecycle support and expert services, from cloud advisory and migration to operations.

Support & Professional Services

Professional Services \ Expert-led services to design, migrate, and optimize your cloud journey Support Plans \ Flexible support for every stage — from startup to enterprise

Partner Support Program \ Priority technical support for partners, with dedicated managers and faster issue resolution

Contact us

Connect With Us \

Talk to a sales expert and get a custom quote for your business

Language

  • English
  • 简体中文
  • 繁體中文
  • 日本語
  • Bahasa Indonesia

Locale

Visit aliyun.com

Documentation

Alibaba Cloud Model Studio

User Guide (Models) User Guide (Application) API Reference (Models) API Reference (Application)

Search for Help Content

Getting Started

The Beginner's Guide

Well-Architected Framework

AI & Machine Learning

Platform For AI

Alibaba Cloud Model Studio

DashVector

Artificial Intelligence Recommendation

OpenSearch

Image Search

Machine Translation

Intelligent Speech Interaction

Optimization Solver

Intelligent Computing LINGJUN

Computing

Elastic Compute Service

Elastic GPU Service

Elastic Container Instance

Dedicated Host

Compute Nest

Simple Application Server

Cloud Box

Auto Scaling

Elastic High Performance Computing

Batch Compute (Deprecated)

Function Compute

Serverless App Engine

ENS

Elastic Desktop Service

App Streaming

WUYING Terminal

Cloud Phone

Edge Network Acceleration

Alibaba Cloud Linux

AgentBay

Container

Container Service for Kubernetes

Container Compute Service

Container Registry

Storage

Object Storage Service

Cloud Parallel File Storage

File Storage NAS

Tablestore

Storage Capacity Unit

Simple Log Service

Cloud Backup

Intelligent Media Management

Drive and Photo Service

Data Transport

Cloud Storage Gateway

Data Online Migration

Hybrid Cloud Storage Array

Storage Services Overview

Backup and Disaster Recovery Center

Networking and CDN

Server Load Balancer

Elastic IP Address

Internet Shared Bandwidth

Data Transfer Plan

Virtual Private Cloud

NAT Gateway

PrivateLink

Alibaba Cloud DNS PrivateZone

Network Intelligence Service

Cloud Data Transfer

IPv6 Gateway

Anycast Elastic IP Address

Cloud Enterprise Network

Global Accelerator

VPN Gateway

Smart Access Gateway

Express Connect

CDN

Edge Security Acceleration

Cloud Network Well-architected Design Guidelines

Security

Anti-DDoS

Web Application Firewall

Cloud Firewall

Security Center

Bastionhost

Secure Access Service Edge

Certificate Management Service

Key Management Service

Data Security Center

Identity as a Service

Fraud Detection

AI Guardrails

Captcha

Blockchain as a Service

ID Verification

Managed Security Service

Middleware

Enterprise Distributed Application Service

Microservices Engine

Alibaba Cloud Service Mesh

SchedulerX

ApsaraMQ for RocketMQ

ApsaraMQ for Kafka

ApsaraMQ for RabbitMQ

ApsaraMQ for MQTT

Simple Message Queue (formerly MNS)

CloudFlow

EventBridge

Application Real-Time Monitoring Service

Managed Service for Prometheus

Managed Service for Grafana

Managed Service for OpenTelemetry

Performance Testing

STAROps

Databases

ApsaraDB Console

PolarDB

ApsaraDB RDS

ApsaraDB for OceanBase (Deprecated)

Tair (Redis® OSS-Compatible)

Lindorm

Time Series Database

ApsaraDB for MongoDB

ApsaraDB for HBase

ApsaraDB for Memcache

ApsaraDB for MyBase

AnalyticDB

ApsaraDB for ClickHouse

ApsaraDB for SelectDB

Data Transmission Service

Database Autonomy Service

Data Management

Database Gateway - Deprecated

ApsaraDB for Cassandra - Deprecated

Analytics Computing

MaxCompute

Hologres

Realtime Compute for Apache Flink

Elasticsearch

Vector Retrieval Service for Milvus

E-MapReduce

Data Lake Formation

DataV

Quick BI

Quick Audience

Quick Tracking

DataWorks

DataHub

Dataphin

Media Services

ApsaraVideo VOD

ApsaraVideo Live

Intelligent Media Services

ApsaraVideo Media Processing

Apsara Video SDK

Enterprise Services & Cloud Communication

Energy Expert

CloudQuotation

Salesforce on Alibaba Cloud

GoChina ICP Filing Assistant

Marketplace

Alibaba Mail

Direct Mail

Short Message Service

Voice Service

Phone Number Verification Service

Cell Phone Number Service

Chat App Message Service

Financial Intelligence Engine

Domain Names and Websites

Domain Names

ICP Filing

Alibaba Cloud DNS

End User Computing

Elastic Desktop Service

App Streaming

WUYING Terminal

Cloud Phone

AgentBay

Internet of Things

IoT Platform

Serverless

Serverless App Engine

CloudFlow

EventBridge

Simple Message Queue (formerly MNS)

Function Compute

Developer Tools

OpenAPI Explorer

Alibaba Cloud SDK

Cloud Shell

Resource Orchestration Service

Alibaba Cloud CLI

BSS OpenAPI

Terraform

Pulumi

Ticket System API

Mobile Platform as a Service

Alibaba Cloud DevOps

API Gateway

Cloud Control API

AI Coding Assistant Lingma

Cloud Skills Portal

Migration & O&M Management

CloudOps Orchestration Service

Cloud Monitor

Intelligent Advisor

Cloud Governance Center

ActionTrail

Cloud Config

Resource Access Management

Resource Management

Cloud Architect Design Tools

Migration Hub

Server Migration Center

Service Catalog

Logic Composer

Quota Center

CloudSSO

HTTPDNS

Solutions

SAP

SuperApp

OpenLake

Membership Service

Expenses and Costs

Account Center

More

Support

Legal

Tech Share Terms and Conditions

After Sales Support

China Gateway Program

Service Level Objectives

Management Console

Security Control

This topic was translated by AI and is currently in queue for revision by our editors. Alibaba Cloud does not guarantee the accuracy of AI-translated content. Request expedited revision

Alibaba Cloud Model Studio provides video generation models for general-purpose creation (text-to-video, image-to-video, reference-to-video, video editing) and vertical scenarios (digital human lip-syncing, image-to-action, video character swapping, emoji creation).

Model overview

Service deployment scope > Compare scopesGlobal > Compute resources for model inference are scheduled globally.International > Compute resources for model inference are scheduled globally, excluding Chinese Mainland.US > Compute resources for model inference are restricted to the US.Chinese Mainland > Compute resources for model inference are restricted to Chinese Mainland.
RegionVirginiaSingaporeVirginiaBeijing
Supported modelsWan - text-to-video Wan - image-to-video - first frame Wan - reference-to-videoWan - text-to-video Wan - image-to-video Wan - image-to-video - first frame Wan - image-to-video - first and last frames Wan - reference-to-video Wan - general video editing Wan - image-to-action Wan - video character swapWan - text-to-video Wan - image-to-video - first frameWan - text-to-video Wan - image-to-video Wan - image-to-video - first frame Wan - image-to-video - first and last frames Wan - reference-to-video Wan - general video editing Wan - digital human Wan - image-to-action Wan - video character swap AnimateAnyone EMO LivePortrait Emoji VideoRetalk Video style transform

Model selection

  • General video generation

    • To generate a video from a text prompt, use Wan - text-to-video.

    • To generate a cinematic shot from a single image, use Wan - image-to-video - first frame.

    • To control the transition between a starting and an ending image, use Wan - image-to-video - first and last frames.

    • To replicate a character's appearance and voice from reference videos to match a new script, use Wan - reference-to-video.

  • Digital human lip-syncing: Animates static photos to speak, sing, or narrate --- background stays fixed while the face, head, and body move.

    • For the most natural results, including facial expressions, head, and body movements, use Wan - digital human. This model replaces EMO.

    • For videos longer than 20 seconds with simple head movements, such as news reports, use LivePortrait.

  • Video motion transfer : This feature keeps the background of the photo static and animates the person using motion from a reference video. Use Wan - image-to-action.

  • Video character swapping : This feature replaces the person in a video with a person from an image while keeping the original background. Use Wan - video character swap.

  • Dance replacement: Replaces the dancer in a video with a person from an image. For best quality, use Wan - image-to-action and Wan - video character swapping. If budget is limited, use AnimateAnyone.

  • Video lip replacement: This feature replaces the lip movements in an existing video to match new audio. Use VideoRetalk.

  • Emoji creation: This feature creates emojis using fixed-style templates. Use Emoji.

  • Video redrawing: To use fixed-style templates, use Video style transform. To describe styles freely using prompts, use Wan - video editing.

  • Video editing: For all the following tasks, use Wan - general video editing.

    • Local video editing: Replace elements such as subjects or clothing, or remove bystanders.

    • Video extension: Extend short videos, for example, from 1 second to 5 seconds.

    • Video frame expansion: Convert landscape videos to portrait mode or fill in missing borders.

    • Multi-image reference generation: Fuse background and subject images to create a video.

Supported models

Wan - text-to-video

Generates videos from text prompts. It supports text and audio input to create cinematic, multi-shot videos.

API reference | Model pricing | Try online: Singapore, Virginia, Beijing

Global

International

US

Chinese mainland

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).

ModelFeaturesInput modalityOutput video specifications
wan2.6-t2v RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, audioResolution options: 720P, 1080P Video duration: 5s, 10s, 15s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.7-t2v RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-t2vVideo with audio Multi-shot narrative, audio-video synchronizationText, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.5-t2v-previewVideo with audio Audio-video synchronizationText, audioResolution options: 480P, 720P, 1080P Video duration: 5s, 10s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.2-t2v-plusSilent video > Improved stability and success rate compared to the 2.1 model.TextResolution options: 480P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.1-t2v-turboSilent videoTextResolution options: 480P, 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.1-t2v-plusSilent videoTextResolution options: 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the US deployment scope, model inference compute resources are restricted to the United States. Static data is stored in your selected region. Supported region: US (Virginia).

ModelFeaturesInput modalityOutput video specifications
wan2.6-t2v-us RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, audioResolution options: 720P, 1080P Video duration: 5s, 10s, 15s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.7-t2v RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-t2vVideo with audio Multi-shot narrative, audio-video synchronizationText, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.5-t2v-previewVideo with audio Audio-video synchronizationText, audioResolution options: 480P, 720P, 1080P Video duration: 5s, 10s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.2-t2v-plusSilent video > Improved stability and success rate compared to the 2.1 model.TextResolution options: 480P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wanx2.1-t2v-turboSilent videoTextResolution options: 480P, 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wanx2.1-t2v-plusSilent videoTextResolution options: 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
Input promptOutput video (wan2.6, multi-shot video)
Input promptOutput video (wan2.6, multi-shot video)
Shot from a low angle, in a medium close-up, with warm tones, mixed lighting (the practical light from the desk lamp blends with the overcast light from the window), side lighting, and a central composition. In a classic detective office, wooden bookshelves are filled with old case files and ashtrays. A green desk lamp illuminates a case file spread out in the center of the desk. A fox, wearing a dark brown trench coat and a light gray fedora, sits in a leather chair, its fur crimson, its tail resting lightly on the edge, its fingers slowly turning yellowed pages. Outside, a steady drizzle falls beneath a blue sky, streaking the glass with meandering streaks. It slowly raises its head, its ears twitching slightly, its amber eyes gazing directly at the camera, its mouth clearly moving as it speaks in a smooth, cynical voice: 'The case was cold, colder than a fish in winter. But every chicken has its secrets, and I, for one, intended to find them '.

Wan - image-to-video

The Wan image-to-video model is upgraded with multimodal input (text/image/audio/video) and supports three tasks: first-frame-to-video, first-and-last-frame-to-video, and video continuation.

API reference | Model pricing | Prompt guide

International

Chinese Mainland

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.7-i2v RecommendedVideo with audio First-frame-to-video, first-and-last-frame-to-video, video continuation, video continuation with last frame control Multi-shot narrative, audio-video synchronizationText, image, audio, videoResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.7-i2v RecommendedVideo with audio First-frame-to-video, first-and-last-frame-to-video, video continuation, video continuation with last frame control Multi-shot narrative, audio-video synchronizationText, image, audio, videoResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)

Wan - image-to-video - first frame

Generates a video from a specified first-frame image. This model accepts text, a first-frame image, and audio as input to generate cinematic, multi-shot videos.

API reference | Model pricing | Try online: Singapore, Virginia, Beijing

Global

International

US

Chinese Mainland

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).

ModelFeaturesInput modalityOutput video specifications
wan2.6-i2v RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, image, audioResolution options: 720P, 1080P Video duration: 5s, 10s, 15s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.6-i2v-flash RecommendedVideo with audio, silent video Multi-shot narrative, audio-video synchronizationText, image, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-i2v RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, image, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.5-i2v-previewVideo with audio Audio-video synchronizationText, image, audioResolution options: 480P, 720P, 1080P Video duration: 5s, 10s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.2-i2v-flashSilent video > 50% faster than the 2.1 model.Text, imageResolution options: 480P, 720P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.2-i2v-plusSilent video > Improved stability and success rate compared to the 2.1 model.Text, imageResolution options: 480P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.1-i2v-plusSilent videoText, imageResolution options: 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.1-i2v-turboSilent videoText, imageResolution options: 480P, 720P Video duration: 3s, 4s, 5s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the US deployment scope, model inference compute resources are restricted to the United States. Static data is stored in your selected region. Supported region: US (Virginia).

ModelFeaturesInput modalityOutput video specifications
wan2.6-i2v-us RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, image, audioResolution options: 720P, 1080P Video duration: 5s, 10s, 15s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.6-i2v-flash RecommendedVideo with audio, silent video Multi-shot narrative, audio-video synchronizationText, image, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-i2v RecommendedVideo with audio Multi-shot narrative, audio-video synchronizationText, image, audioResolution options: 720P, 1080P Video duration: [2s, 15s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.5-i2v-previewVideo with audio Audio-video synchronizationText, image, audioResolution options: 480P, 720P, 1080P Video duration: 5s, 10s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.2-i2v-flashSilent video > 50% faster than the 2.1 model.Text, imageResolution options: 480P, 720P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.2-i2v-plusSilent video > Improved stability and success rate compared to the 2.1 model.Text, imageResolution options: 480P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wanx2.1-i2v-plusSilent videoText, imageResolution options: 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wanx2.1-i2v-turboSilent videoText, imageResolution options: 480P, 720P Video duration: 3s, 4s, 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
Input promptInput first frame image and audioOutput video (wan2.6, multi-shot video)
Input promptInput first frame image and audioOutput video (wan2.6, multi-shot video)
An urban fantasy art scene. A dynamic graffiti art character. A teenager made of spray paint comes to life from a concrete wall. He performs an English rap at high speed while striking a classic, energetic rapper pose. The scene is set under an urban railway bridge at night. The lighting comes from a single streetlight, creating a cinematic atmosphere with high energy and amazing detail. The audio of the video consists entirely of his rap, with no other dialogue or noise.Input audio:

Wan - image-to-video - first and last frames

Generates a video that smoothly transitions between specified first and last frame images. This model accepts text, first and last frame images, and audio as input to generate cinematic, multi-shot videos.

API reference | Model pricing | Try online

International

Chinese Mainland

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.2-kf2v-flash RecommendedSilent video > Improved stability and success rate compared to the 2.1 model.Text, imageResolution options: 480P, 720P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.1-kf2v-plusSilent videoText, imageResolution options: 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.2-kf2v-flash RecommendedSilent video > Improved stability and success rate compared to the 2.1 model.Text, imageResolution options: 480P, 720P, 1080P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
wanx2.1-kf2v-plusSilent videoText, imageResolution options: 720P Video duration: 5s Defined specifications: 30 fps, MP4 (H.264 encoding)
Input first frame imageInput last frame imageInput promptOutput video
Input first frame imageInput last frame imageInput promptOutput video
Realistic style. A small black cat looks up at the sky curiously. The camera starts at eye level, gradually rises, and ends with a top-down shot of the cat's curious gaze.

Wan - reference-to-video

Make characters from a specified video perform actions. Input a video and a text prompt to generate an output video that maintains character consistency.

API reference | Model pricing

Global

International

Chinese mainland

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).

ModelFeaturesInput modalityOutput video specifications
wan2.6-r2v RecommendedVideo with audio Single-role/multi-role video generation Multi-shot narrative, audio-video synchronizationText, videoResolution options: 720P, 1080P Video duration: 5s, 10s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.7-r2v RecommendedVideo with audio Multi-entity reference-to-video; supports configuring voice timbre for each entity.Text, image, video, audioResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-r2v-flashVideo with audio, silent video Single-role/multi-role video generation Multi-shot narrative, audio-video synchronization > Faster generation, cost-effective.Text, image, videoResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-r2vVideo with audio Single-role/multi-role video generation Multi-shot narrative, audio-video synchronizationText, image, videoResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.7-r2v RecommendedVideo with audio Multi-entity reference-to-video lets you configure voice timbre for each entity.Text, image, video, audioResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-r2v-flashVideo with audio, silent video Single-role/multi-role video generation Multi-shot narrative, audio-video synchronization > Faster generation, cost-effective.Text, image, videoResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.6-r2vVideo with audio Single-role/multi-role video generation Multi-shot narrative, audio-video synchronizationText, image, videoResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
Input reference video 1 (role: little girl)Input reference video 2 (role: alarm clock)Input promptOutput video (multi-role dialogue)
Input reference video 1 (role: little girl)Input reference video 2 (role: alarm clock)Input promptOutput video (multi-role dialogue)
character1 says to character2: “I’ll rely on you tomorrow morning!” character2 replies: “You can count on me!”

Wan - video editing

Video editing model. Accepts text, image, and video multimodal input to perform various video generation and editing tasks.

Video editing 2.7 API reference | Video editing 2.1 API reference | Model pricing

International

Chinese Mainland

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.7-videoedit RecommendedVideo with audio, silent video (depends on the input video) Instruction-based editing, video migrationText, image, videoResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wan2.1-vace-plusSilent video Multi-image reference, video redrawing, local editing, video extension, video frame extensionText, image, videoResolution options: 720P Video duration: Up to 5s Defined specifications: 30 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.7-videoedit RecommendedVideo with audio, silent video (depends on the input video) Instruction-based editing, video migrationText, image, videoResolution options: 720P, 1080P Video duration: [2s, 10s] (integer) Defined specifications: 30 fps, MP4 (H.264 encoding)
wanx2.1-vace-plusSilent video Multi-image reference, video redrawing, local editing, video extension, video frame extensionText, image, videoResolution options: 720P Video duration: Up to 5s Defined specifications: 30 fps, MP4 (H.264 encoding)

Video editing 2.1

  • Feature 1: Multi-image reference
Input reference image 1 (reference entity)Input reference image 2 (reference background)Input promptOutput video
Video shows a girl gracefully walking out from the depths of an ancient, misty forest. Her steps are light, and the camera captures her every nimble moment. When she stops and looks around at the lush woods, a smile of surprise and joy blossoms on her face. This moment, frozen in the interplay of light and shadow, records her wonderful encounter with nature.
  • Feature 2: Video redrawing
Input videoInput promptOutput video
The video shows a black steampunk-style car, driven by a gentleman, adorned with gears and copper pipes. The background is a steam-powered candy factory with retro elements, creating a vintage and playful scene.
  • Feature 3: Local video editing
Input videoInput mask image (the white area indicates the editing area)Input promptOutput video
The video shows a Parisian-style French cafe where a lion in a suit elegantly sips coffee. It holds a coffee cup in one hand, taking a gentle sip with a contented expression. The cafe is tastefully decorated, with soft hues and warm lighting illuminating the area where the lion is.
  • Feature 4: Video extension
Input first video segment (1s)Input promptOutput video (extended video is 5s)
A dog wearing sunglasses skateboards on a street, 3D cartoon.
  • Feature 5: Video frame extension
Input videoInput promptOutput video
An elegant woman passionately plays the violin, with a full symphony orchestra behind her.

Wan - digital human

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Digital human lip-syncing animates a person or cartoon character in an image to speak, sing, narrate, or perform. You provide an image and an audio file, and the model automatically generates a video with synchronized lip movements, facial expressions, and head and body motions.

Image detection API reference | Video generation API reference | Model pricing

ModelFeaturesInput modalityOutput description
ModelFeaturesInput modalityOutput description
wan2.2-s2v-detectImage detectionImageOutput detection status: Pass or Fail
wan2.2-s2vVideo generation Video with audioImage, audioResolution options: 480P, 720P Video duration: Up to 20s (follows audio duration) Defined specifications: - 480P: 16 fps, MP4 (H.264 encoding) - 720P: 30 fps, MP4 (H.264 encoding)
Input example (character image + audio)Output video (lip-sync)
Input example (character image + audio)Output video (lip-sync)
Input audio:

Wan - image to action

Animates a person from an image using motion from a reference video. You provide an image and a video, and the model generates a video that applies the motion from the reference video to the person, while keeping the background of the original image static.

API reference | Model pricing

International

Chinese mainland

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.2-animate-moveVideo with audio, silent video (depends on the input video) - Standard mode wan-std: Fast generation, cost-effective. - Professional mode wan-pro: More realistic results.Image, videoResolution options: 720P Video duration: 2s < duration < 30s Defined specifications: - Standard mode wan-std: 15 fps, MP4 (H.264 encoding) - Professional mode wan-pro: 25 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.2-animate-moveVideo with audio, silent video (depends on the input video) - Standard mode wan-std: Fast generation, cost-effective. - Professional mode wan-pro: More realistic results.Image, videoResolution options: 720P Video duration: 2s < duration < 30s Defined specifications: - Standard mode wan-std: 15 fps, MP4 (H.264 encoding) - Professional mode wan-pro: 25 fps, MP4 (H.264 encoding)
Input character imageInput reference videoOutput video (standard modewan-std )Output video (professional modewan-pro )
Input character imageInput reference videoOutput video (standard modewan-std )Output video (professional modewan-pro )

Wan - video character swap

Replaces a character in a video with one from a reference image. You provide a source video and a reference image, and the model generates an output video that retains the original background. This feature is ideal for use cases like face swapping and full character replacement.

API reference | Model pricing

International

Chinese mainland

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

ModelFeaturesInput modalityOutput video specifications
wan2.2-animate-mixVideo with audio, silent video (depends on the input video) - Standard mode wan-std: Fast generation, cost-effective. - Professional mode wan-pro: More realistic results.Image, videoResolution options: 720P Video duration: 2s < duration < 30s Defined specifications: - Standard mode wan-std: 15 fps, MP4 (H.264 encoding) - Professional mode wan-pro: 25 fps, MP4 (H.264 encoding)

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

ModelFeaturesInput modalityOutput video specifications
wan2.2-animate-mixVideo with audio, silent video (depends on the input video) - Standard mode wan-std: Fast generation, cost-effective. - Professional mode wan-pro: More realistic results.Image, videoResolution options: 720P Video duration: 2s < duration < 30s Defined specifications: - Standard mode wan-std: 15 fps, MP4 (H.264 encoding) - Professional mode wan-pro: 25 fps, MP4 (H.264 encoding)
Input videoInput character image for replacementOutput video (standard modewan-std )Output video (professional modewan-pro )
Input videoInput character image for replacementOutput video (standard modewan-std )Output video (professional modewan-pro )

AnimateAnyone

Note

  • Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • We recommend using Wan - image-to-action and Wan - video character swapping instead of AnimateAnyone. These models offer better quality, while AnimateAnyone is a more cost-effective option.

Designed specifically for dancing, this model replaces the dancer in a video with a person from an image. You provide an image and a video to generate an output video in one of two ways: 1. Retain the image background. 2. Retain the video background.

Image detection API reference | Action template generation API reference | Video generation API reference | Model pricing

ModelFeaturesInput modalityOutput description
ModelFeaturesInput modalityOutput description
animate-anyone-detect-gen2Image detectionImageOutput detection status: Pass or Fail
animate-anyone-template-gen2Dance video template generation > Extracts an action template from a dance video.VideoOutputs a dance action template ID.
animate-anyone-gen2Video generation Silent videoImage, video, dance action template IDVideo resolution options: 720P Video duration: 2s ≤ duration ≤ 60s Defined specifications: 15 fps, MP4 (H.264 encoding)
Input character imageInput dance videoOutput video (generated with image background)Output video (generated with video background)
Input character imageInput dance videoOutput video (generated with image background)Output video (generated with video background)

EMO

Note

  • Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • Consider using Wan - digital human as an alternative to EMO. Wan - digital human provides better results, while EMO is a more cost-effective option.

Generates singing and performance videos from an image. You provide an image and an audio file, and the model automatically generates a video with synchronized lip movements, facial expressions, and head motions.

Image detection API reference | Video generation API reference | Model pricing

ModelFeaturesInput modalityOutput description
ModelFeaturesInput modalityOutput description
emo-detect-v1Image detectionImageOutput detection status: Pass or Fail
emo-v1Video generation Video with audioImage, audioVideo resolution: - 1:1 aspect ratio: Fixed at 512 × 512 - 3:4 aspect ratio: Fixed at 512 × 704 Video duration: Up to 60s Defined specifications: 15 fps, MP4 (H.264 encoding)
Input example (portrait image + audio)Output video (lip-sync singing)
Input example (portrait image + audio)Output video (lip-sync singing)
Input audio:

LivePortrait

Note

  • Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • Consider using Wan - digital human as an alternative to LivePortrait. Wan - digital human delivers higher quality results, while LivePortrait is a more cost-effective option. Note that LivePortrait is suitable for generating long videos (over 20 seconds).

Generates narration videos from an image by animating the person in the image to deliver news or tell stories. You provide an Image and an Audio file, and the model automatically generates a video with synchronized lip movements, facial expressions, and slight head motions.

Image detection API reference | Video generation API reference | Model pricing

ModelFeaturesInput modalityOutput description
ModelFeaturesInput modalityOutput description
liveportrait-detectImage detectionImageOutput detection status: Pass or Fail
liveportraitVideo generation Video with audioImage, audioVideo resolution: Follows the input image, up to nearly 4K (4096 × 4096). Video duration: 1s < duration < 180s Video frame rate: 15 fps ≤ frame rate ≤ 30 fps Video format: MP4 (H.264 encoding)
Input example (portrait image + audio)Output video (lip-sync voiceover)
Input example (portrait image + audio)Output video (lip-sync voiceover)
Input audio:

Emoji

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Creates emojis using fixed emoji templates. You provide an image and an emoji template ID to generate an emoji video.

Image detection API reference |  Video generation API reference |  Model pricing

ModelFeaturesInput modalityOutput description
ModelFeaturesInput modalityOutput description
emoji-detect-v1Image detectionImageOutput detection status: Pass or Fail
emoji-v1Video generation Silent videoImage, emoji template IDVideo resolution: Fixed at 512 × 512 Video duration: Up to 5s (follows template duration) Defined specifications: 15 fps, MP4 (H.264 encoding)
Input portrait imageOutput video ("disgusted" emoji)
Input portrait imageOutput video ("disgusted" emoji)

VideoRetalk

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Lip sync: Replaces the lip movements in a video to match a new audio track. You provide a video and an audio file, and the model generates an output video with synchronized lip movements.

API reference |  Model pricing

ModelFeaturesInput modalityOutput video specifications
ModelFeaturesInput modalityOutput video specifications
videoretalkVideo with audioVideo, audioVideo resolution: Follows the input video, up to nearly 2K (2048 × 2048). Video duration: 2s < duration < 120s Video frame rate: 15 fps ≤ frame rate ≤ 60 fps Video format: MP4 (H.264 encoding)
Input example (character broadcast video + audio)Output video (lip-sync replacement)
Input example (character broadcast video + audio)Output video (lip-sync replacement)
Input audio:

Video style transform

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Applies a new artistic style to a video based on a predefined style template. You provide a video and a style transfer ID to generate a restyled video.

API reference |  Model pricing

ModelFeaturesInput modalityOutput video specifications
ModelFeaturesInput modalityOutput video specifications
video-style-transformVideo with audio, silent video > Depends on the input video.Video redraw style IDVideo resolution: Follows the input video, up to nearly 4K (4096 × 4096). Video duration: Up to 30s Video frame rate: 15 fps ≤ frame rate ≤ 25 fps Video format: MP4 (H.264 encoding)
Input videoOutput video (style transfer: "Japanese manga")
Input videoOutput video (style transfer: "Japanese manga")

Previous:NoneNext: Product introduction

Is this page helpful?

Model overview

Model selection

Supported models

Wan - text-to-video

Wan - image-to-video

Wan - image-to-video - first frame

Wan - image-to-video - first and last frames

Wan - reference-to-video

Wan - video editing

Wan - digital human

Wan - image to action

Wan - video character swap

AnimateAnyone

EMO

LivePortrait

Emoji

VideoRetalk

Video style transform

Contact Us

Sales Support

Live-chat with our sales team or get in touch with a business development professional in your region.

Contact Sales

Technical Support

Open a ticket and get quick help from our technical team.

Open a Ticket >

Connect & Report Abuse

We look forward to your suggestion.

Post a Suggestion > Report Abuse >

Chat now with Alibaba Cloud Customer Service to assist you in finding the right products and services to meet your needs.

\ \ Hi, I'm Alibaba Cloud AI Assistant!\ \ I can help with questions and solutions.

Why Alibaba Cloud

About Alibaba Cloud

Asia Accelerator

Our Global Network

Global Offices

Trust Center

Case Studies

Analyst Reports

Products & Pricings

Pricing Calculator

ECS

SAS

Model Studio

Database

Security

SMS

Solutions

Financial Services

Retail Services

Media Services

Gaming Services

ISV Solutions

Engage

Developer Community

Partner Network

Startups

Marketplace

Join Alibaba Cloud

Resources & Support

Developer Learning Hub

Documentation Center

Training & Certification

Service Notices

Submit a Ticket

Security Report

Qwen Cloud

Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links

© 2009-2026 Copyright by Alibaba Cloud All rights reserved

© 2009-2026 Copyright by Alibaba Cloud All rights reserved

Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links