Appearance
Voice cloning
Reference, synced 2026-06-13.
flowchart TD n0["Models"] n1["Overview"] n2["Products"] n3["Solutions"] n4["Pricing"] n5["Resources"] n6["Partners"] n7["Support"] n8["Language"] n0 --> n1 n1 --> n2 n2 --> n3 n3 --> n4 n4 --> n5 n5 --> n6 n6 --> n7 n7 --> n8
Free access Accelerate Delivery with Fixed-Cost Agentic CodingWatch how it works
Models
Empowering AI innovation for both enterprises and developers with Alibaba Cloud’s best-in-class Qwen models, AI-native apps, and AI solutions.
Alibaba Cloud Model Studio \ Enterprise-grade large model service and application development platform.
Try Visual Model \ Supports image understanding, image generation, and video generation.
Models
HappyHorse-1.0-T2V \ Cinematic creative generation, ultimate dynamic details Qwen3-VL-Plus \ Native VL, spatial reasoning, 1M-context video analysis Wan2.7-VideoEdit \ Supports both localized and global editing with prompt
Qwen3.6-Plus \ Native multimodal, 1M context, agentic coding Wan2.7-Image-Pro \ Interactive editing, long-text rendering, precise prompt following Qwen-Plus \ Balanced intelligence, efficient inference, production-ready performance
Qwen-Image-2.0 \ Professional infographics, exquisite photorealism Z-Image-Turbo \ Ultra-fast image generation, high throughput, cost-optimized inference Qwen3-Coder-Next \ Multi-turn tool interactions, future-ready development support
Wan2.7-T2V \ High-fidelity T2V, 15s duration, advanced camera control Wan2.7-I2V \ Cinematic I2V with emotional depth and visceral impact Wan2.7-R2V \ Up to 5 mixed image/video inputs and audio timbre cloning
GenAI Application
Qoder \ Intelligent coding assistant, available for enterprise-dedicated deployment. Qoder CN \ AI-powered coding assistant that boosts developer productivity with intelligent code completion, AI chat, multi-file editing, and task automation.
AI Service
Model Experience \ Experience full-scale, multimodal model capabilities online. Platform for AI \ An AI-native algorithm engineering platform for end-to-end modeling, training, and inference service deployment. Fine-tune Video Generation Model \ Customize Wan’s text-to-video capabilities through model fine-tuning to meet your unique requirements.
AI Use Case
AI Savings Plan Hot \ Save up to 47% on AI costs. Limited-time offer tailored to your usage. AI Video Creation \ Elevate your professional video production with Wan 2.6.
AI Token Plan \ One plan. Multiple models. Big Savings with a Fixed Subscription. AI Image Creation \ All-in-one creative suite for copywriting, image generation, and poster design.
Overview
As a global full-stack AI leader, Alibaba Cloud aims to make computing accessible to everyone and help worldwide customers accelerate innovation.
Why Alibaba Cloud
About Alibaba Cloud \ AI Powered Cloud Technology Our Global Network \ Explore our global presence and deployment regions around the world Our Global Offices \ With offices in 4 continents, we're always close to where it matters.
Asia Accelerator \ Accelerate Success in Asia with Alibaba Cloud Go Global \ Benefits of our Global Alliance Trust Center \ Empowering enterprises with a secure, compliant, and globally trusted cloud infrastructure
Customers and Insights
Olympic Games \ Alibaba Cloud Powers Olympic Games with AI-powered cloud technology Case Studies \ Learn how customers are scaling their businesses on Alibaba Cloud Analyst Reports \ Learn what the top industry analyst firms are saying about Alibaba Cloud
What's New
Events and Webinars \ Quick access to upcoming and on-demand events Product Updates \ Stay informed of the latest innovations Press Room \ Latest news and media releases
Products
Featured ProductsAI & Machine Learning Computing Container Storage Networking & CDN Security Middleware Database Analytics ComputingMedia ServicesEnterprise Services & Cloud CommunicationDomain Names and WebsitesEnd User ComputingServerlessDeveloper ToolsMigration & O&M ManagementApsara Stack
Featured Products
Alibaba Cloud Model Studio \ Supercharge your AI journey effortlessly with industry-leading GenAI models ApsaraDB RDS \ Store and manage your business data, with automated monitoring and backups Certificate Management Service (Original SSL Certificate) \ Create a safe and secure connection between your website and users
Elastic Compute Service (ECS) \ Host websites anywhere and scale enterprise workloads Container Service for Kubernetes (ACK) \ Run and scale containerized applications on managed Kubernetes infrastructure Object Storage Service (OSS) \ Store large amounts of data in the cloud and access it anywhere, anytime
Simple Application Server (SAS) \ All-in-one services for fast deployment Elastic IP Address (EIP) \ Manage your public IPs independently to improve internet network quality Domain Names and Website \ Get the perfect domain name to suit your every need
Related Programs
Solutions
Solutions by Industry Technical Solutions AI WebsitesNetworking Security and ComplianceData and AnalyticsEnterprise Service and ApplicationCloud MigrationCloud NativeHybrid CloudSMB solutions
Financial Services \ Innovate faster with Alibaba Cloud Games \ Grow your game rapidly with high global availability
New Retail \ Alibaba Cloud enables digital retail transformation to fuel growth and realize an omnichannel customer experience throughout the consumer journey. Media and Entertainment \ Ready your content for today's media market with a digitalized media journey
Supply Chain \ Power your supply chain with intelligent, efficient, and reliable solutions Sports \ Digitizing the sports industry with intelligent tech
Sustainability \ Achieve a sustainable future with low-carbon and energy-efficient technologies
Pricing
Flexible options like pay-as-you-go and clear billing rules to meet diverse business needs.
Overview & Tools
Pricing Calculator \ Get an instant pricing estimate based on your usage and needs Free Trial \ Try our 80+ cloud products for free.
Pricing Options \ Get the most out of Alibaba Cloud with flexible pricing
Optimize your cost
Migrate & Save \ Superior Performance At Lower Pricing. Save up to 50%. Promotion Center \ Unlock the latest Alibaba Cloud offers & promos
Resources
Official documentation, extensive tools, training resources, and a community to grow and innovate in the cloud.
Technical Resources
Documentation \ Product guides and FAQs Architecture Center \ Design reliable, secure, and efficient cloud architecture. Intelligent Solution Explorer \ Find the right solution for you, powered by AI
Blog \ Latest cloud insights and developer trends Whitepapers \ Research that explores the how and why behind our technology
Training&Certification
Alibaba Cloud Academy \ Build cloud skills and earn certifications with expert-led training.
Developer Hub
Alibaba Cloud Project Hub \ Explore real-world projects built by developers using our platform. Our Developer MVPs \ Celebrating the developers who lead, build, and inspire our community
Partners
Partner-first strategy offering collaborative product, sales, and service models, plus high-quality partner solutions that complement Alibaba Cloud’s capabilities.
Marketplace
AI Alliance for ISVs \ Partner with us to build and grow AI solutions together ISV Benefits \ Unlock resources, market access, and go-to-market support as an ISV partner
Alibaba Cloud Marketplace \ Explore ready-to-deploy solutions from our partners and ISVs
Find a Partner
Partner Hub \ Find your ideal partner in no time
Become a Partner
Partner Network \ A partner portal for Alibaba Cloud Channel, Technology, MSP partner and other partner programs
Support
Full-lifecycle support and expert services, from cloud advisory and migration to operations.
Support & Professional Services
Professional Services \ Expert-led services to design, migrate, and optimize your cloud journey Support Plans \ Flexible support for every stage — from startup to enterprise
Partner Support Program \ Priority technical support for partners, with dedicated managers and faster issue resolution
Contact us
Connect With Us \
Talk to a sales expert and get a custom quote for your business
Language
- English
- 简体中文
- 繁體中文
- 日本語
- Bahasa Indonesia
Locale
Visit aliyun.com
Documentation
Alibaba Cloud Model Studio
User Guide (Models) User Guide (Application) API Reference (Models) API Reference (Application)
Search for Help Content
Getting Started
The Beginner's Guide
Well-Architected Framework
AI & Machine Learning
Platform For AI
Alibaba Cloud Model Studio
DashVector
Artificial Intelligence Recommendation
OpenSearch
Image Search
Machine Translation
Intelligent Speech Interaction
Optimization Solver
Intelligent Computing LINGJUN
Computing
Elastic Compute Service
Elastic GPU Service
Elastic Container Instance
Dedicated Host
Compute Nest
Simple Application Server
Cloud Box
Auto Scaling
Elastic High Performance Computing
Batch Compute (Deprecated)
Function Compute
Serverless App Engine
ENS
Elastic Desktop Service
App Streaming
WUYING Terminal
Cloud Phone
Edge Network Acceleration
Alibaba Cloud Linux
AgentBay
Container
Container Service for Kubernetes
Container Compute Service
Container Registry
Storage
Object Storage Service
Cloud Parallel File Storage
File Storage NAS
Tablestore
Storage Capacity Unit
Simple Log Service
Cloud Backup
Intelligent Media Management
Drive and Photo Service
Data Transport
Cloud Storage Gateway
Data Online Migration
Hybrid Cloud Storage Array
Storage Services Overview
Backup and Disaster Recovery Center
Networking and CDN
Server Load Balancer
Elastic IP Address
Internet Shared Bandwidth
Data Transfer Plan
Virtual Private Cloud
NAT Gateway
PrivateLink
Alibaba Cloud DNS PrivateZone
Network Intelligence Service
Cloud Data Transfer
IPv6 Gateway
Anycast Elastic IP Address
Cloud Enterprise Network
Global Accelerator
VPN Gateway
Smart Access Gateway
Express Connect
CDN
Edge Security Acceleration
Cloud Network Well-architected Design Guidelines
Security
Anti-DDoS
Web Application Firewall
Cloud Firewall
Security Center
Bastionhost
Secure Access Service Edge
Certificate Management Service
Key Management Service
Data Security Center
Identity as a Service
Fraud Detection
AI Guardrails
Captcha
Blockchain as a Service
ID Verification
Managed Security Service
Middleware
Enterprise Distributed Application Service
Microservices Engine
Alibaba Cloud Service Mesh
SchedulerX
ApsaraMQ for RocketMQ
ApsaraMQ for Kafka
ApsaraMQ for RabbitMQ
ApsaraMQ for MQTT
Simple Message Queue (formerly MNS)
CloudFlow
EventBridge
Application Real-Time Monitoring Service
Managed Service for Prometheus
Managed Service for Grafana
Managed Service for OpenTelemetry
Performance Testing
STAROps
Databases
ApsaraDB Console
PolarDB
ApsaraDB RDS
ApsaraDB for OceanBase (Deprecated)
Tair (Redis® OSS-Compatible)
Lindorm
Time Series Database
ApsaraDB for MongoDB
ApsaraDB for HBase
ApsaraDB for Memcache
ApsaraDB for MyBase
AnalyticDB
ApsaraDB for ClickHouse
ApsaraDB for SelectDB
Data Transmission Service
Database Autonomy Service
Data Management
Database Gateway - Deprecated
ApsaraDB for Cassandra - Deprecated
Analytics Computing
MaxCompute
Hologres
Realtime Compute for Apache Flink
Elasticsearch
Vector Retrieval Service for Milvus
E-MapReduce
Data Lake Formation
DataV
Quick BI
Quick Audience
Quick Tracking
DataWorks
DataHub
Dataphin
Media Services
ApsaraVideo VOD
ApsaraVideo Live
Intelligent Media Services
ApsaraVideo Media Processing
Apsara Video SDK
Enterprise Services & Cloud Communication
Energy Expert
CloudQuotation
Salesforce on Alibaba Cloud
GoChina ICP Filing Assistant
Marketplace
Alibaba Mail
Direct Mail
Short Message Service
Voice Service
Phone Number Verification Service
Cell Phone Number Service
Chat App Message Service
Financial Intelligence Engine
Domain Names and Websites
Domain Names
ICP Filing
Alibaba Cloud DNS
End User Computing
Elastic Desktop Service
App Streaming
WUYING Terminal
Cloud Phone
AgentBay
Internet of Things
IoT Platform
Serverless
Serverless App Engine
CloudFlow
EventBridge
Simple Message Queue (formerly MNS)
Function Compute
Developer Tools
OpenAPI Explorer
Alibaba Cloud SDK
Cloud Shell
Resource Orchestration Service
Alibaba Cloud CLI
BSS OpenAPI
Terraform
Pulumi
Ticket System API
Mobile Platform as a Service
Alibaba Cloud DevOps
API Gateway
Cloud Control API
AI Coding Assistant Lingma
Cloud Skills Portal
Migration & O&M Management
CloudOps Orchestration Service
Cloud Monitor
Intelligent Advisor
Cloud Governance Center
ActionTrail
Cloud Config
Resource Access Management
Resource Management
Cloud Architect Design Tools
Migration Hub
Server Migration Center
Service Catalog
Logic Composer
Quota Center
CloudSSO
HTTPDNS
Solutions
SAP
SuperApp
OpenLake
Membership Service
Expenses and Costs
Account Center
More
Support
Legal
Tech Share Terms and Conditions
After Sales Support
China Gateway Program
Service Level Objectives
Management Console
Security Control
Voice Cloning creates a highly realistic custom voice from a 10- to 20-second audio sample, with no model training required.
Overview
Voice Cloning lets you build personalized voice assistants, branded audio broadcasts, and custom narration.
Model Studio supports Voice Cloning through the following model families:
CosyVoice: Create voices through the DashScope SDK or HTTP API. Supports real-time speech synthesis. Available in the China (Beijing) and Singapore regions.
Qwen-TTS: Create voices through the HTTP API. Supports both real-time and non-real-time speech synthesis. Available in the China (Beijing) and Singapore regions.
For a detailed comparison and guidance on choosing a model family, see Speech synthesis.
Prerequisites
Configure an API key and set it as an environment variable.
If you call the API through the DashScope SDK, install the latest SDK.
Prepare an audio file that meets the Audio requirements.
Quick start
Voice cloning involves three steps:
Prepare the audio: Prepare an audio file that meets the Audio requirements.
Create a voice: Call the Voice Cloning API to upload the audio and create a voice. In the
target_modelparameter, specify the speech synthesis model to bind to the voice.Synthesize speech: Call the speech synthesis API and pass the voice ID returned when you created the voice.
CosyVoice voice cloning
Important
CosyVoice voice cloning is available in the China (Beijing) region (v3.5, v3, v2, and v1 series) and the Singapore region (v3 series only).
Step 1: Create a voice
Call the Voice Cloning API to upload an audio file and create a voice. The url parameter is the accessible URL of the audio file; prefix sets a prefix for the voice name.
China (Beijing) region URL. The URL varies by region.
Singapore region URL. Replace WorkspaceId with your actual workspace ID. The URL varies by region.
curl
curl -X POST 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/customization' \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "voice-enrollment",
"input": {
"action": "create_voice",
"target_model": "cosyvoice-v3-plus",
"prefix": "myvoice",
"url": "https://your-audio-url.wav"
}
}'Step 2: Synthesize speech with the cloned voice
Replace voice_id in the following code with the value returned in the previous step.
python
# coding=utf-8
import dashscope
from dashscope.audio.tts_v2 import *
import os
# The API keys for the Singapore and Beijing regions are different. To obtain an API key, visit: https://www.alibabacloud.com/help/en/model-studio/get-api-key
# If the environment variable is not set, replace the following line with your Model Studio API key: dashscope.api_key = "sk-xxx"
dashscope.api_key = os.environ.get('DASHSCOPE_API_KEY')
# Singapore region URL. Replace WorkspaceId with your actual workspace ID. The URL varies by region.
dashscope.base_websocket_api_url='wss://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api-ws/v1/inference'
# Use the same model for both voice cloning and speech synthesis
model = "cosyvoice-v3-plus"
# Replace the voice parameter with the custom voice generated by cloning
voice = "voice_id"
# Create a SpeechSynthesizer instance with the model and voice parameters
synthesizer = SpeechSynthesizer(model=model, voice=voice)
# Send text for synthesis and get binary audio
audio = synthesizer.call("How is the weather today?")
# The first call incurs extra latency for establishing the WebSocket connection
print('[Metric] requestId: {}, first packet latency: {} ms'.format(
synthesizer.get_last_request_id(),
synthesizer.get_first_package_delay()))
# Save the audio to a local file
with open('output.mp3', 'wb') as f:
f.write(audio)Qwen-TTS voice cloning
The example uses the local audio file voice.mp3. Before running the code, replace voice.mp3 with the path to your audio file.
Important
The target_model set during voice creation must exactly match the model used for speech synthesis. Otherwise, synthesis fails.
Python
cURL
python
import os
import requests
import base64
import pathlib
import dashscope
# ======= Constants =======
DEFAULT_TARGET_MODEL = "qwen3-tts-vc-2026-01-22" # Use the same model for both voice cloning and speech synthesis
DEFAULT_PREFERRED_NAME = "guanyu"
DEFAULT_AUDIO_MIME_TYPE = "audio/mpeg"
VOICE_FILE_PATH = "voice.mp3" # Relative path to the local audio file used for voice cloning
def create_voice(file_path: str,
target_model: str = DEFAULT_TARGET_MODEL,
preferred_name: str = DEFAULT_PREFERRED_NAME,
audio_mime_type: str = DEFAULT_AUDIO_MIME_TYPE) -> str:
"""
Create a custom voice and return the voice parameter.
"""
# The API keys for the Singapore and Beijing regions are different. To obtain an API key, visit: https://www.alibabacloud.com/help/en/model-studio/get-api-key
# If the environment variable is not set, replace the following line with your Model Studio API key: api_key = "sk-xxx"
api_key = os.getenv("DASHSCOPE_API_KEY")
file_path_obj = pathlib.Path(file_path)
if not file_path_obj.exists():
raise FileNotFoundError(f"Audio file not found: {file_path}")
base64_str = base64.b64encode(file_path_obj.read_bytes()).decode()
data_uri = f"data:{audio_mime_type};base64,{base64_str}"
# Singapore region URL. Replace WorkspaceId with your actual workspace ID. The URL varies by region.
url = "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/customization"
payload = {
"model": "qwen-voice-enrollment", # Do not modify this value
"input": {
"action": "create",
"target_model": target_model,
"preferred_name": preferred_name,
"audio": {"data": data_uri}
}
}
headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json"
}
resp = requests.post(url, json=payload, headers=headers)
if resp.status_code != 200:
raise RuntimeError(f"Failed to create voice: {resp.status_code}, {resp.text}")
try:
return resp.json()["output"]["voice"]
except (KeyError, ValueError) as e:
raise RuntimeError(f"Failed to parse voice response: {e}")
if __name__ == '__main__':
# Singapore region URL. Replace WorkspaceId with your actual workspace ID. The URL varies by region.
dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'
text = "How is the weather today?"
response = dashscope.MultiModalConversation.call(
model=DEFAULT_TARGET_MODEL,
api_key=os.getenv("DASHSCOPE_API_KEY"),
text=text,
voice=create_voice(VOICE_FILE_PATH), # Replace the voice parameter with the custom voice generated by cloning
stream=False
)
print(response)Replace data with the actual path to your audio file.
China (Beijing) region URL. The URL varies by region.
Step 1: Create a voice
Singapore region URL. Replace WorkspaceId with your actual workspace ID. The URL varies by region.
curl
curl -X POST 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/customization' \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen-voice-enrollment",
"input": {
"action": "create",
"target_model": "qwen3-tts-vc-2026-01-22",
"preferred_name": "guanyu",
"audio": {
"data": "https://xxx.wav"
}
}
}'Step 2: Synthesize speech with the cloned voice
Replace YOUR_VOICE_ID with the voice value from the previous step's response.
China (Beijing) region URL. The URL varies by region.
Singapore region URL. Replace WorkspaceId with your actual workspace ID. The URL varies by region.
curl
curl -X POST 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen3-tts-vc-2026-01-22",
"input": {
"text": "How is the weather today?",
"voice": "YOUR_VOICE_ID"
}
}'Audio requirements
The quality of the input audio directly affects the cloning result. Each model family has different audio requirements. Prepare your audio sample according to the requirements of your target model.
CosyVoice
MiniMax
Qwen-TTS
| Item | Requirement |
| Supported formats | WAV (16-bit), MP3, M4A |
| Duration | 10 to 20 seconds recommended. Maximum 60 seconds. |
| File size | 10 MB or less |
| Sample rate | 16 kHz or higher |
| Channels | Mono or stereo. For stereo audio, only the first channel is processed. Make sure the first channel contains valid speech. |
| Content | The audio must contain at least 5 seconds of continuous, clear speech. Brief pauses in the remaining portion must not exceed 2 seconds. Avoid background music, ambient noise, or other voices. Use normal-speed spoken audio; don't upload songs or singing. |
| Supported languages | Varies by the speech synthesis model specified through the target_model parameter: - cosyvoice-v2: Chinese (Mandarin), English - cosyvoice-v3-flash: Chinese (Mandarin, Cantonese, Northeastern, Gansu, Guizhou, Henan, Hubei, Jiangxi, Minnan, Ningxia, Shanxi, Shaanxi, Shandong, Shanghainese, Sichuanese, Tianjin, Yunnan), English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, Vietnamese - cosyvoice-v3-plus: Chinese (Mandarin, Cantonese, Northeastern, Gansu, Guizhou, Henan, Hubei, Jiangxi, Minnan, Ningxia, Shanxi, Shaanxi, Shandong, Shanghainese, Sichuanese, Tianjin, Yunnan), English, French, German, Japanese, Korean, Russian - cosyvoice-v3.5-plus, cosyvoice-v3.5-flash: Chinese (Mandarin, Cantonese, Henan, Hubei, Minnan, Ningxia, Shaanxi, Shandong, Shanghainese, Sichuanese), English, French, German, Japanese, Korean, Russian, Portuguese, Thai, Indonesian, Vietnamese |
| Item | Requirement |
| Supported formats | MP3, M4A, WAV |
| Duration | At least 10 seconds. Maximum 5 minutes. |
| File size | 20 MB or less |
| Content | The audio must contain continuous, clear speech with no background sound. Pauses must not exceed 2 seconds. Avoid background music, ambient noise, or other voices throughout the recording. Use normal-speed spoken audio. Don't upload songs or singing recordings. |
| Supported languages | No restrictions |
| Item | Requirement |
| Supported formats | WAV (16-bit), MP3, M4A |
| Duration | 10 to 20 seconds recommended. Maximum 60 seconds. |
| File size | Less than 10 MB |
| Sample rate | 24 kHz or higher |
| Channels | Mono |
| Content | The audio must contain at least 3 seconds of continuous, clear speech. Brief pauses in the remaining portion must not exceed 2 seconds. Avoid background music, ambient noise, or other voices. Use normal-speed spoken audio; don't upload songs or singing. |
| Supported languages | Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian |
Note
For best cloning results, follow the Recording tips when preparing your audio sample.
Recording tips
High-quality input audio produces better cloning results.
Recording equipment
Use a smartphone, digital voice recorder, or professional recording device. For best results, use a device with a sample rate of 24 kHz or higher.
Recording environment
Venue
Record in a small, enclosed space of 10 square meters or less.
Prefer a room with sound-absorbing materials such as acoustic foam, carpet, or curtains.
Avoid open halls, conference rooms, classrooms, and other spaces with high reverberation.
Noise control
Outdoor noise: Close doors and windows to block traffic, construction, and other external sounds.
Indoor noise: Turn off air conditioners, fans, fluorescent light ballasts, and other appliances. To identify hidden noise sources, record a few seconds of ambient sound and play it back at higher volume.
Reverberation control
Reverberation blurs the sound and reduces clarity.
Reduce reflections from smooth surfaces: close curtains, open closet doors, and drape clothing or blankets over desks and cabinets.
Use irregularly shaped objects such as bookshelves and upholstered furniture to diffuse sound.
Recording script
No specific content restrictions. Match the script to the target use case when possible.
Avoid short phrases such as "Hello" or "Yes." Use complete sentences.
Keep the content coherent and avoid frequent pauses. Aim for at least 3 seconds of continuous speech without interruption.
Maintain a consistent pace throughout the recording. Speaking too fast at the start or finish may cause stuttering in the synthesized speech.
Include natural emotional expression — warmth, friendliness, or seriousness. Avoid robotic delivery.
Don't include sensitive content such as political, sexual, or violent material. This causes the cloning request to fail.
Recording workflow
The following example uses a typical bedroom as the recording space. Complete the noise reduction and reverberation control steps described above, then:
Review the script, decide on a tone and persona, then record naturally.
Hold the recording device about 10 cm from your mouth to avoid plosive distortion or a weak signal.
Manage custom voices
After creating a voice with Qwen-TTS or CosyVoice, you can query and manage your voices through the API.
List voices: Get a list of all custom voices under your account.
Get voice details: View details of a specific voice, such as the creation time and the bound speech synthesis model.
Delete voices: Delete custom voices you no longer need to free up quota.
For API endpoints and parameter details, see API reference.
Quota and billing
Voice quota and automatic cleanup
Total voice limit: Each Alibaba Cloud Model Studio account has a separate limit of 1,000 custom voices for CosyVoice and 1,000 for Qwen-TTS. The two quotas are counted independently.
Automatic cleanup: If a voice isn't used in any speech synthesis request for one year, the system automatically deletes it.
Billing rules
CosyVoice: Voice creation is free.
Qwen-TTS: Each voice creation costs USD 0.01. Failed creations aren't charged.
Free quota (Singapore region only):
You get 1,000 free voice creations during the first 90 days after activating Alibaba Cloud Model Studio.
Failed creations don't consume the free quota.
Deleting a voice doesn't restore the free quota.
After the free quota is used up or the 90-day window expires, voice creation is billed at USD 0.01 per voice.
Supported scope
Available models vary by region:
Singapore
China (Beijing)
Use a Singapore-region API Key when calling the following models:
CosyVoice: cosyvoice-v3-plus, cosyvoice-v3-flash
Qwen-TTS:
Qwen3-TTS-VC-Realtime : qwen3-tts-vc-realtime-2026-01-15 (latest snapshot), qwen3-tts-vc-realtime-2025-11-27 (snapshot)
Qwen3-TTS-VC: qwen3-tts-vc-2026-01-22 (latest snapshot)
Use a China (Beijing)-region API Key when calling the following models:
CosyVoice: cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-plus, cosyvoice-v3-flash, cosyvoice-v2
Qwen-TTS:
Qwen3-TTS-VC-Realtime : qwen3-tts-vc-realtime-2026-01-15 (latest snapshot), qwen3-tts-vc-realtime-2025-11-27 (snapshot)
Qwen3-TTS-VC: qwen3-tts-vc-2026-01-22 (latest snapshot)
API reference
Voice Cloning
FAQ
Q: Can I use a created voice with different speech synthesis models?
No. A voice is bound to a specific speech synthesis model through the target_model parameter during voice creation and can't be used across models. To use the same audio recording with multiple models, create a separate voice for each model.
Q: How long does a cloned voice remain valid?
Voices created with Qwen-TTS and CosyVoice are valid indefinitely by default. If a voice goes unused for one year, the system automatically deletes it. For details, see Voice quota and automatic cleanup. Save your voice IDs and use the query API to check whether a voice is still available.
Q: Does poor audio quality affect the cloning result?
Yes. The quality of the input audio directly affects the cloning result. Background noise, reverberation, and overlapping voices all reduce the similarity and naturalness of the cloned voice. Follow the Audio requirements and Recording tips when preparing your audio sample.
Previous: Non-real-time speech synthesisNext: Voice Design
Is this page helpful?
Overview
Prerequisites
Quick start
CosyVoice voice cloning
Qwen-TTS voice cloning
Audio requirements
Recording tips
Recording equipment
Recording environment
Recording script
Recording workflow
Manage custom voices
Quota and billing
Voice quota and automatic cleanup
Billing rules
Supported scope
API reference
FAQ
Q: Can I use a created voice with different speech synthesis models?
Q: How long does a cloned voice remain valid?
Q: Does poor audio quality affect the cloning result?
Contact Us
Sales Support
Live-chat with our sales team or get in touch with a business development professional in your region.
Contact Sales
Technical Support
Open a ticket and get quick help from our technical team.
Open a Ticket >
Connect & Report Abuse
We look forward to your suggestion.
Post a Suggestion > Report Abuse >
\ \ Hi, I'm Alibaba Cloud AI Assistant!\ \ I can help with questions and solutions.
Why Alibaba Cloud
About Alibaba Cloud
Asia Accelerator
Our Global Network
Global Offices
Trust Center
Case Studies
Analyst Reports
Products & Pricings
Pricing Calculator
ECS
SAS
Model Studio
Database
Security
SMS
Solutions
Financial Services
Retail Services
Media Services
Gaming Services
ISV Solutions
Engage
Developer Community
Partner Network
Startups
Marketplace
Join Alibaba Cloud
Resources & Support
Developer Learning Hub
Documentation Center
Training & Certification
Service Notices
Submit a Ticket
Security Report
Qwen Cloud
Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links
© 2009-2026 Copyright by Alibaba Cloud All rights reserved
- YouTube
- TikTok
- contact.us@alibabacloud.com
- Call Us Now
- Discord
© 2009-2026 Copyright by Alibaba Cloud All rights reserved
Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links