Skip to content

Qwen-OCR API reference

Reference, synced 2026-06-13.

flowchart TD
  n0["Models"]
  n1["Overview"]
  n2["Products"]
  n3["Solutions"]
  n4["Pricing"]
  n5["Resources"]
  n6["Partners"]
  n7["Support"]
  n8["Language"]
  n0 --> n1
  n1 --> n2
  n2 --> n3
  n3 --> n4
  n4 --> n5
  n5 --> n6
  n6 --> n7
  n7 --> n8

Free access Accelerate Delivery with Fixed-Cost Agentic CodingWatch how it works

Models

Empowering AI innovation for both enterprises and developers with Alibaba Cloud’s best-in-class Qwen models, AI-native apps, and AI solutions.

Alibaba Cloud Model Studio \ Enterprise-grade large model service and application development platform.

Try Visual Model \ Supports image understanding, image generation, and video generation.

Models

HappyHorse-1.0-T2V \ Cinematic creative generation, ultimate dynamic details Qwen3-VL-Plus \ Native VL, spatial reasoning, 1M-context video analysis Wan2.7-VideoEdit \ Supports both localized and global editing with prompt

Qwen3.6-Plus \ Native multimodal, 1M context, agentic coding Wan2.7-Image-Pro \ Interactive editing, long-text rendering, precise prompt following Qwen-Plus \ Balanced intelligence, efficient inference, production-ready performance

Qwen-Image-2.0 \ Professional infographics, exquisite photorealism Z-Image-Turbo \ Ultra-fast image generation, high throughput, cost-optimized inference Qwen3-Coder-Next \ Multi-turn tool interactions, future-ready development support

Wan2.7-T2V \ High-fidelity T2V, 15s duration, advanced camera control Wan2.7-I2V \ Cinematic I2V with emotional depth and visceral impact Wan2.7-R2V \ Up to 5 mixed image/video inputs and audio timbre cloning

GenAI Application

Qoder \ Intelligent coding assistant, available for enterprise-dedicated deployment. Qoder CN \ AI-powered coding assistant that boosts developer productivity with intelligent code completion, AI chat, multi-file editing, and task automation.

AI Service

Model Experience \ Experience full-scale, multimodal model capabilities online. Platform for AI \ An AI-native algorithm engineering platform for end-to-end modeling, training, and inference service deployment. Fine-tune Video Generation Model \ Customize Wan’s text-to-video capabilities through model fine-tuning to meet your unique requirements.

AI Use Case

AI Savings Plan Hot \ Save up to 47% on AI costs. Limited-time offer tailored to your usage. AI Video Creation \ Elevate your professional video production with Wan 2.6.

AI Token Plan \ One plan. Multiple models. Big Savings with a Fixed Subscription. AI Image Creation \ All-in-one creative suite for copywriting, image generation, and poster design.

Overview

As a global full-stack AI leader, Alibaba Cloud aims to make computing accessible to everyone and help worldwide customers accelerate innovation.

Why Alibaba Cloud

About Alibaba Cloud \ AI Powered Cloud Technology Our Global Network \ Explore our global presence and deployment regions around the world Our Global Offices \ With offices in 4 continents, we're always close to where it matters.

Asia Accelerator \ Accelerate Success in Asia with Alibaba Cloud Go Global \ Benefits of our Global Alliance Trust Center \ Empowering enterprises with a secure, compliant, and globally trusted cloud infrastructure

Customers and Insights

Olympic Games \ Alibaba Cloud Powers Olympic Games with AI-powered cloud technology Case Studies \ Learn how customers are scaling their businesses on Alibaba Cloud Analyst Reports \ Learn what the top industry analyst firms are saying about Alibaba Cloud

What's New

Events and Webinars \ Quick access to upcoming and on-demand events Product Updates \ Stay informed of the latest innovations Press Room \ Latest news and media releases

Products

Featured ProductsAI & Machine Learning Computing Container Storage Networking & CDN Security Middleware Database Analytics ComputingMedia ServicesEnterprise Services & Cloud CommunicationDomain Names and WebsitesEnd User ComputingServerlessDeveloper ToolsMigration & O&M ManagementApsara Stack

Alibaba Cloud Model Studio \ Supercharge your AI journey effortlessly with industry-leading GenAI models ApsaraDB RDS \ Store and manage your business data, with automated monitoring and backups Certificate Management Service (Original SSL Certificate) \ Create a safe and secure connection between your website and users

Elastic Compute Service (ECS) \ Host websites anywhere and scale enterprise workloads Container Service for Kubernetes (ACK) \ Run and scale containerized applications on managed Kubernetes infrastructure Object Storage Service (OSS) \ Store large amounts of data in the cloud and access it anywhere, anytime

Simple Application Server (SAS) \ All-in-one services for fast deployment Elastic IP Address (EIP) \ Manage your public IPs independently to improve internet network quality Domain Names and Website \ Get the perfect domain name to suit your every need

Solutions

Solutions by Industry Technical Solutions AI WebsitesNetworking Security and ComplianceData and AnalyticsEnterprise Service and ApplicationCloud MigrationCloud NativeHybrid CloudSMB solutions

Financial Services \ Innovate faster with Alibaba Cloud Games \ Grow your game rapidly with high global availability

New Retail \ Alibaba Cloud enables digital retail transformation to fuel growth and realize an omnichannel customer experience throughout the consumer journey. Media and Entertainment \ Ready your content for today's media market with a digitalized media journey

Supply Chain \ Power your supply chain with intelligent, efficient, and reliable solutions Sports \ Digitizing the sports industry with intelligent tech

Sustainability \ Achieve a sustainable future with low-carbon and energy-efficient technologies

Pricing

Flexible options like pay-as-you-go and clear billing rules to meet diverse business needs.

Overview & Tools

Pricing Calculator \ Get an instant pricing estimate based on your usage and needs Free Trial \ Try our 80+ cloud products for free.

Pricing Options \ Get the most out of Alibaba Cloud with flexible pricing

Optimize your cost

Migrate & Save \ Superior Performance At Lower Pricing. Save up to 50%. Promotion Center \ Unlock the latest Alibaba Cloud offers & promos

Resources

Official documentation, extensive tools, training resources, and a community to grow and innovate in the cloud.

Technical Resources

Documentation \ Product guides and FAQs Architecture Center \ Design reliable, secure, and efficient cloud architecture. Intelligent Solution Explorer \ Find the right solution for you, powered by AI

Blog \ Latest cloud insights and developer trends Whitepapers \ Research that explores the how and why behind our technology

Training&Certification

Alibaba Cloud Academy \ Build cloud skills and earn certifications with expert-led training.

Developer Hub

Alibaba Cloud Project Hub \ Explore real-world projects built by developers using our platform. Our Developer MVPs \ Celebrating the developers who lead, build, and inspire our community

Partners

Partner-first strategy offering collaborative product, sales, and service models, plus high-quality partner solutions that complement Alibaba Cloud’s capabilities.

Marketplace

AI Alliance for ISVs \ Partner with us to build and grow AI solutions together ISV Benefits \ Unlock resources, market access, and go-to-market support as an ISV partner

Alibaba Cloud Marketplace \ Explore ready-to-deploy solutions from our partners and ISVs

Find a Partner

Partner Hub \ Find your ideal partner in no time

Become a Partner

Partner Network \ A partner portal for Alibaba Cloud Channel, Technology, MSP partner and other partner programs

Support

Full-lifecycle support and expert services, from cloud advisory and migration to operations.

Support & Professional Services

Professional Services \ Expert-led services to design, migrate, and optimize your cloud journey Support Plans \ Flexible support for every stage — from startup to enterprise

Partner Support Program \ Priority technical support for partners, with dedicated managers and faster issue resolution

Contact us

Connect With Us \

Talk to a sales expert and get a custom quote for your business

Language

  • English
  • 简体中文
  • 繁體中文
  • 日本語
  • Bahasa Indonesia

Locale

Visit aliyun.com

Documentation

Alibaba Cloud Model Studio

User Guide (Models) User Guide (Application) API Reference (Models) API Reference (Application)

Search for Help Content

Getting Started

The Beginner's Guide

Well-Architected Framework

AI & Machine Learning

Platform For AI

Alibaba Cloud Model Studio

DashVector

Artificial Intelligence Recommendation

OpenSearch

Image Search

Machine Translation

Intelligent Speech Interaction

Optimization Solver

Intelligent Computing LINGJUN

Computing

Elastic Compute Service

Elastic GPU Service

Elastic Container Instance

Dedicated Host

Compute Nest

Simple Application Server

Cloud Box

Auto Scaling

Elastic High Performance Computing

Batch Compute (Deprecated)

Function Compute

Serverless App Engine

ENS

Elastic Desktop Service

App Streaming

WUYING Terminal

Cloud Phone

Edge Network Acceleration

Alibaba Cloud Linux

AgentBay

Container

Container Service for Kubernetes

Container Compute Service

Container Registry

Storage

Object Storage Service

Cloud Parallel File Storage

File Storage NAS

Tablestore

Storage Capacity Unit

Simple Log Service

Cloud Backup

Intelligent Media Management

Drive and Photo Service

Data Transport

Cloud Storage Gateway

Data Online Migration

Hybrid Cloud Storage Array

Storage Services Overview

Backup and Disaster Recovery Center

Networking and CDN

Server Load Balancer

Elastic IP Address

Internet Shared Bandwidth

Data Transfer Plan

Virtual Private Cloud

NAT Gateway

PrivateLink

Alibaba Cloud DNS PrivateZone

Network Intelligence Service

Cloud Data Transfer

IPv6 Gateway

Anycast Elastic IP Address

Cloud Enterprise Network

Global Accelerator

VPN Gateway

Smart Access Gateway

Express Connect

CDN

Edge Security Acceleration

Cloud Network Well-architected Design Guidelines

Security

Anti-DDoS

Web Application Firewall

Cloud Firewall

Security Center

Bastionhost

Secure Access Service Edge

Certificate Management Service

Key Management Service

Data Security Center

Identity as a Service

Fraud Detection

AI Guardrails

Captcha

Blockchain as a Service

ID Verification

Managed Security Service

Middleware

Enterprise Distributed Application Service

Microservices Engine

Alibaba Cloud Service Mesh

SchedulerX

ApsaraMQ for RocketMQ

ApsaraMQ for Kafka

ApsaraMQ for RabbitMQ

ApsaraMQ for MQTT

Simple Message Queue (formerly MNS)

CloudFlow

EventBridge

Application Real-Time Monitoring Service

Managed Service for Prometheus

Managed Service for Grafana

Managed Service for OpenTelemetry

Performance Testing

STAROps

Databases

ApsaraDB Console

PolarDB

ApsaraDB RDS

ApsaraDB for OceanBase (Deprecated)

Tair (Redis® OSS-Compatible)

Lindorm

Time Series Database

ApsaraDB for MongoDB

ApsaraDB for HBase

ApsaraDB for Memcache

ApsaraDB for MyBase

AnalyticDB

ApsaraDB for ClickHouse

ApsaraDB for SelectDB

Data Transmission Service

Database Autonomy Service

Data Management

Database Gateway - Deprecated

ApsaraDB for Cassandra - Deprecated

Analytics Computing

MaxCompute

Hologres

Realtime Compute for Apache Flink

Elasticsearch

Vector Retrieval Service for Milvus

E-MapReduce

Data Lake Formation

DataV

Quick BI

Quick Audience

Quick Tracking

DataWorks

DataHub

Dataphin

Media Services

ApsaraVideo VOD

ApsaraVideo Live

Intelligent Media Services

ApsaraVideo Media Processing

Apsara Video SDK

Enterprise Services & Cloud Communication

Energy Expert

CloudQuotation

Salesforce on Alibaba Cloud

GoChina ICP Filing Assistant

Marketplace

Alibaba Mail

Direct Mail

Short Message Service

Voice Service

Phone Number Verification Service

Cell Phone Number Service

Chat App Message Service

Financial Intelligence Engine

Domain Names and Websites

Domain Names

ICP Filing

Alibaba Cloud DNS

End User Computing

Elastic Desktop Service

App Streaming

WUYING Terminal

Cloud Phone

AgentBay

Internet of Things

IoT Platform

Serverless

Serverless App Engine

CloudFlow

EventBridge

Simple Message Queue (formerly MNS)

Function Compute

Developer Tools

OpenAPI Explorer

Alibaba Cloud SDK

Cloud Shell

Resource Orchestration Service

Alibaba Cloud CLI

BSS OpenAPI

Terraform

Pulumi

Ticket System API

Mobile Platform as a Service

Alibaba Cloud DevOps

API Gateway

Cloud Control API

AI Coding Assistant Lingma

Cloud Skills Portal

Migration & O&M Management

CloudOps Orchestration Service

Cloud Monitor

Intelligent Advisor

Cloud Governance Center

ActionTrail

Cloud Config

Resource Access Management

Resource Management

Cloud Architect Design Tools

Migration Hub

Server Migration Center

Service Catalog

Logic Composer

Quota Center

CloudSSO

HTTPDNS

Solutions

SAP

SuperApp

OpenLake

Membership Service

Expenses and Costs

Account Center

More

Support

Legal

Tech Share Terms and Conditions

After Sales Support

China Gateway Program

Service Level Objectives

Management Console

Security Control

Extract text, structured data, and key information from images using the Qwen-OCR model. Qwen-OCR supports two API protocols: the OpenAI-compatible API and the DashScope API.

For use cases and getting-started guidance, see Text extraction (Qwen-OCR).

OpenAI-compatible API

Endpoints

RegionSDKbase_urlHTTP endpoint
RegionSDKbase_urlHTTP endpoint
Singaporehttps://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions
US (Virginia)https://dashscope-us.aliyuncs.com/compatible-mode/v1POST https://dashscope-us.aliyuncs.com/compatible-mode/v1/chat/completions
China (Beijing)https://dashscope.aliyuncs.com/compatible-mode/v1POST https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions

Prerequisites

Get an API key and set it as an environment variable. If you use the OpenAI SDK, install the SDK.

Quick start

Use the OpenAI-compatible chat completions endpoint. Send a user message with an image URL and text prompt. The model extracts text and returns it in choices[0].message.content.

Non-streaming

Python

python
from openai import OpenAI
import os

PROMPT_TICKET_EXTRACTION = """
Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image.
You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?).
Return the data in JSON format as follows: {'invoice_number': 'xxx', 'departure_station': 'xxx', 'arrival_station': 'xxx', 'departure_date_and_time':'xxx', 'seat_number': 'xxx','ticket_price':'xxx', 'id_card_number': 'xxx', 'passenger_name': 'xxx'},
"""

try:
    client = OpenAI(
        # If the environment variable is not configured, replace with: api_key="sk-xxx"
        api_key=os.getenv("DASHSCOPE_API_KEY"),
        # Singapore region. For US (Virginia), use https://dashscope-us.aliyuncs.com/compatible-mode/v1
        # For China (Beijing), use https://dashscope.aliyuncs.com/compatible-mode/v1
        base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
    )
    completion = client.chat.completions.create(
        model="qwen-vl-ocr-2025-11-20",
        messages=[\
            {\
                "role": "user",\
                "content": [\
                    {\
                        "type": "image_url",\
                        "image_url": {"url":"https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg"},\
                        # Minimum pixel count. Images below this are upscaled.\
                        "min_pixels": 32 * 32 * 3,\
                        # Maximum pixel count. Images above this are downscaled.\
                        "max_pixels": 32 * 32 * 8192\
                    },\
                    # Custom prompt. Without this, the model uses: "Please output only the text content from the image without any additional descriptions or formatting."\
                    {"type": "text",\
                     "text": PROMPT_TICKET_EXTRACTION}\
                ]\
            }\
        ])
    print(completion.choices[0].message.content)
except Exception as e:
    print(f"Error message: {e}")

Node.js

javascript
import OpenAI from 'openai';

const PROMPT_TICKET_EXTRACTION = `
Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image.
You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?).
Return the data in JSON format as follows: {'invoice_number': 'xxx', 'departure_station': 'xxx', 'arrival_station': 'xxx', 'departure_date_and_time':'xxx', 'seat_number': 'xxx','ticket_price':'xxx', 'id_card_number': 'xxx', 'passenger_name': 'xxx'}
`;

const client = new OpenAI({
  // If the environment variable is not configured, replace with: apiKey: "sk-xxx"
  apiKey: process.env.DASHSCOPE_API_KEY,
  // For China (Beijing), use https://dashscope.aliyuncs.com/compatible-mode/v1
  baseURL: 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1',
});

async function main() {
  const response = await client.chat.completions.create({
    model: 'qwen-vl-ocr-2025-11-20',
    messages: [\
      {\
        role: 'user',\
        content: [\
          { type: 'text', text: PROMPT_TICKET_EXTRACTION},\
          {\
            type: 'image_url',\
            image_url: {\
              url: 'https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg',\
            },\
              // Minimum pixel count. Images below this are upscaled.\
              "min_pixels": 32 * 32 * 3,\
              // Maximum pixel count. Images above this are downscaled.\
              "max_pixels": 32 * 32 * 8192\
          }\
        ]\
      }\
    ],
  });
  console.log(response.choices[0].message.content)
}

main();

curl

bash
curl -X POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "model": "qwen-vl-ocr-2025-11-20",
  "messages": [\
        {\
            "role": "user",\
            "content": [\
                {\
                    "type": "image_url",\
                    "image_url": {"url":"https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg"},\
                    "min_pixels": 3072,\
                    "max_pixels": 8388608\
                },\
                {"type": "text", "text": "Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image. You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?). Return the data in JSON format as follows: {\'invoice_number\': \'xxx\', \'departure_station\': \'xxx\', \'arrival_station\': \'xxx\', \'departure_date_and_time\':\'xxx\', \'seat_number\': \'xxx\',\'ticket_price\':\'xxx\', \'id_card_number\': \'xxx\', \'passenger_name\': \'xxx\'}"}\
            ]\
        }\
    ]
}'

Streaming

Set stream to true to receive results incrementally as the model generates them.

Python

python
import os
from openai import OpenAI

PROMPT_TICKET_EXTRACTION = """
Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image.
You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?).
Return the data in JSON format as follows: {'invoice_number': 'xxx','departure_station': 'xxx', 'arrival_station': 'xxx', 'departure_date_and_time':'xxx', 'seat_number': 'xxx','ticket_price':'xxx', 'id_card_number': 'xxx', 'passenger_name': 'xxx'},
"""

client = OpenAI(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)
completion = client.chat.completions.create(
    model="qwen-vl-ocr-2025-11-20",
    messages=[\
        {\
            "role": "user",\
            "content": [\
                {\
                    "type": "image_url",\
                    "image_url": {"url":"https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg"},\
                    "min_pixels": 32 * 32 * 3,\
                    "max_pixels": 32 * 32 * 8192\
                },\
                {"type": "text","text": PROMPT_TICKET_EXTRACTION}\
            ]\
        }\
    ],
    stream=True,
    stream_options={"include_usage": True}
)

for chunk in completion:
    print(chunk.model_dump_json())

Node.js

javascript
import OpenAI from 'openai';

const PROMPT_TICKET_EXTRACTION = `
Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image.
You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?).
Return the data in JSON format as follows: {'invoice_number': 'xxx', 'departure_station': 'xxx', 'arrival_station': 'xxx', 'departure_date_and_time':'xxx', 'seat_number': 'xxx','ticket_price':'xxx', 'id_card_number': 'xxx', 'passenger_name': 'xxx'}
`;

const openai = new OpenAI({
  apiKey: process.env.DASHSCOPE_API_KEY,
  baseURL: 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1',
});

async function main() {
  const response = await openai.chat.completions.create({
    model: 'qwen-vl-ocr-2025-11-20',
    messages: [\
      {\
        role: 'user',\
        content: [\
          { type: 'text', text: PROMPT_TICKET_EXTRACTION},\
          {\
            type: 'image_url',\
            image_url: {\
              url: 'https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg',\
            },\
              "min_pixels": 32 * 32 * 3,\
              "max_pixels": 32 * 32 * 8192\
          }\
        ]\
      }\
    ],
    stream: true,
    stream_options:{"include_usage": true}
  });
  let fullContent = ""
  console.log("Streaming output content:")
  for await (const chunk of response) {
    if (chunk.choices[0] && chunk.choices[0].delta.content != null) {
      fullContent += chunk.choices[0].delta.content;
      console.log(chunk.choices[0].delta.content);
    }
  }
  console.log(`Full output content: ${fullContent}`)
}

main();

curl

bash
curl -X POST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
  "model": "qwen-vl-ocr-2025-11-20",
  "messages": [\
        {\
            "role": "user",\
            "content": [\
                {\
                    "type": "image_url",\
                    "image_url": {"url":"https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg"},\
                    "min_pixels": 3072,\
                    "max_pixels": 8388608\
                },\
                {"type": "text", "text": "Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image. You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?). Return the data in JSON format as follows: {\'invoice_number\': \'xxx\', \'departure_station\': \'xxx\', \'arrival_station\': \'xxx\', \'departure_date_and_time\':\'xxx\', \'seat_number\': \'xxx\',\'ticket_price\':\'xxx\', \'id_card_number\': \'xxx\', \'passenger_name\': \'xxx\'}"}\
            ]\
        }\
    ],
    "stream": true,
    "stream_options": {"include_usage": true}
}'

Request parameters

ParameterTypeRequiredDescription
ParameterTypeRequiredDescription
modelstringYesModel name. See Recommended models for supported models.
messagesarrayYesAn array of message objects that provides context to the model.

Message object

Each message requires a role (must be user) and a content array with these element types:

ParameterTypeRequiredDescription
ParameterTypeRequiredDescription
typestringYestext for text input, image_url for image input.
textstringNoThe text prompt. Default: "Please output only the text content from the image without any additional descriptions or formatting".
image_url.urlstringYes (when type is image_url)URL or Base64-encoded Data URL of the image. For local files, see Text extraction.
min_pixelsintegerNoMinimum pixel threshold. Images below this value are upscaled. See Image resolution control.
max_pixelsintegerNoMaximum pixel threshold. Images above this value are downscaled. See Image resolution control.

Generation parameters

ParameterTypeDefaultDescription
ParameterTypeDefaultDescription
streambooleanfalseSet to true to receive incremental responses as the model generates output.
stream_options.include_usagebooleanfalseWhen stream is true, set this to true to include token usage in the last chunk.
max_tokensintegerVariesMaximum tokens in the output. Exceeding this truncates the response. See Output token limits.
temperaturefloat0.01Controls output diversity. Higher values produce more varied text. Range: [0, 2).
top_pfloat0.001Nucleus sampling threshold. Higher values increase diversity. Range: (0, 1.0]. Set either temperature or top_p, not both.
top_kinteger1Limits the candidate token set during sampling. If the value is None or greater than 100, the top_k policy is not enabled, and only the top_p policy takes effect. Must be >= 0. Not a standard OpenAI parameter -- pass via extra_body in the Python SDK: extra_body={"top_k": xxx}. In the Node.js SDK or HTTP, pass at the top level.
repetition_penaltyfloat1.0Penalty for repeated sequences. Values above 1.0 reduce repetition. Not a standard OpenAI parameter -- pass via extra_body in the Python SDK.
presence_penaltyfloat0.0Controls content repetition. Range: [-2.0, 2.0]. Positive values reduce repetition.
seedinteger--Ensures reproducible results when the same value is used with identical parameters. Range: [0, 2^31 - 1].
logprobsbooleanfalseSet to true to return log probabilities of output tokens.
top_logprobsinteger0Number of most likely tokens to return per step. Range: [0, 5]. Only effective when logprobs is true.
stopstring or array--Stop words or token IDs. Generation stops when a specified string or token_id appears. Do not mix strings and token_ids in the same array.

Response

Non-streaming response (chat.completion)

json
{
  "id": "chatcmpl-ba21fa91-dcd6-4dad-90cc-6d49c3c39094",
  "choices": [\
    {\
      "finish_reason": "stop",\
      "index": 0,\
      "logprobs": null,\
      "message": {\
        "content": "```json\n{\n    \"seller_name\": \"null\",\n    \"buyer_name\": \"Cai Yingshi\",\n    \"price_excluding_tax\": \"230769.23\",\n    \"organization_code\": \"null\",\n    \"invoice_code\": \"142011726001\"\n}\n```",\
        "refusal": null,\
        "role": "assistant",\
        "annotations": null,\
        "audio": null,\
        "function_call": null,\
        "tool_calls": null\
      }\
    }\
  ],
  "created": 1763283287,
  "model": "qwen-vl-ocr-latest",
  "object": "chat.completion",
  "service_tier": null,
  "system_fingerprint": null,
  "usage": {
    "completion_tokens": 72,
    "prompt_tokens": 1185,
    "total_tokens": 1257,
    "completion_tokens_details": {
      "accepted_prediction_tokens": null,
      "audio_tokens": null,
      "reasoning_tokens": null,
      "rejected_prediction_tokens": null,
      "text_tokens": 72
    },
    "prompt_tokens_details": {
      "audio_tokens": null,
      "cached_tokens": null,
      "image_tokens": 1001,
      "text_tokens": 184
    }
  }
}
FieldTypeDescription
FieldTypeDescription
idstringUnique request identifier.
choicesarrayModel-generated content.
choices[].finish_reasonstringstop when generation completed normally, length when truncated due to token limit.
choices[].indexintegerPosition in the choices array.
choices[].message.contentstringExtracted text or structured output from the model.
choices[].message.rolestringAlways assistant.
choices[].message.refusalstringAlways null.
choices[].message.audioobjectAlways null.
choices[].message.function_callobjectAlways null.
choices[].message.tool_callsarrayAlways null.
createdintegerUNIX timestamp of the request.
modelstringModel used.
objectstringAlways chat.completion.
service_tierstringAlways null.
system_fingerprintstringAlways null.
usage.completion_tokensintegerOutput token count.
usage.prompt_tokensintegerInput token count.
usage.total_tokensintegerSum of prompt_tokens and completion_tokens.
usage.completion_tokens_details.text_tokensintegerText output tokens. Other fields in completion_tokens_details are always null.
usage.prompt_tokens_details.image_tokensintegerImage input tokens.
usage.prompt_tokens_details.text_tokensintegerText input tokens. Other fields in prompt_tokens_details are always null.

Streaming response (chat.completion.chunk)

When stream is true, the response is delivered as a series of Server-Sent Event (SSE) chunks. Each chunk follows the same structure as the non-streaming response, with these differences:

  • object is always chat.completion.chunk.

  • choices[].delta replaces choices[].message. The delta object has the same fields as message.

  • choices[].delta.role is returned only in the first chunk.

  • finish_reason is null during generation, stop on completion, or length if truncated.

  • When include_usage is true, the last chunk has an empty choices array and includes the usage object.

json
{"id":"chatcmpl-f6fbdc0d-78d6-418f-856f-f099c2e4859b","choices":[{"delta":{"content":"","function_call":null,"refusal":null,"role":"assistant","tool_calls":null},"finish_reason":null,"index":0,"logprobs":null}],"created":1764139204,"model":"qwen-vl-ocr-latest","object":"chat.completion.chunk","service_tier":null,"system_fingerprint":null,"usage":null}
{"id":"chatcmpl-f6fbdc0d-78d6-418f-856f-f099c2e4859b","choices":[{"delta":{"content":"```","function_call":null,"refusal":null,"role":null,"tool_calls":null},"finish_reason":null,"index":0,"logprobs":null}],"created":1764139204,"model":"qwen-vl-ocr-latest","object":"chat.completion.chunk","service_tier":null,"system_fingerprint":null,"usage":null}
{"id":"chatcmpl-f6fbdc0d-78d6-418f-856f-f099c2e4859b","choices":[{"delta":{"content":"json","function_call":null,"refusal":null,"role":null,"tool_calls":null},"finish_reason":null,"index":0,"logprobs":null}],"created":1764139204,"model":"qwen-vl-ocr-latest","object":"chat.completion.chunk","service_tier":null,"system_fingerprint":null,"usage":null}
......
{"id":"chatcmpl-f6fbdc0d-78d6-418f-856f-f099c2e4859b","choices":[{"delta":{"content":"","function_call":null,"refusal":null,"role":null,"tool_calls":null},"finish_reason":"stop","index":0,"logprobs":null}],"created":1764139204,"model":"qwen-vl-ocr-latest","object":"chat.completion.chunk","service_tier":null,"system_fingerprint":null,"usage":null}
{"id":"chatcmpl-f6fbdc0d-78d6-418f-856f-f099c2e4859b","choices":[],"created":1764139204,"model":"qwen-vl-ocr-latest","object":"chat.completion.chunk","service_tier":null,"system_fingerprint":null,"usage":{"completion_tokens":141,"prompt_tokens":513,"total_tokens":654,"completion_tokens_details":{"accepted_prediction_tokens":null,"audio_tokens":null,"reasoning_tokens":null,"rejected_prediction_tokens":null,"text_tokens":141},"prompt_tokens_details":{"audio_tokens":null,"cached_tokens":null,"image_tokens":332,"text_tokens":181}}}

Image resolution control

min_pixels and max_pixels control image resizing before processing. Token-to-pixel ratio depends on model version:

ModelPixels per tokenmin_pixelsdefault (minimum)max_pixelsdefaultmax_pixelsmaximum
ModelPixels per tokenmin_pixelsdefault (minimum)max_pixelsdefaultmax_pixelsmaximum
qwen-vl-ocr-latest, qwen-vl-ocr-2025-11-2032 x 32 = 1,0243,072 (3 tokens)8,388,608 (8,192 tokens)30,720,000 (30,000 tokens)
qwen-vl-ocr, qwen-vl-ocr-2025-08-28, and earlier28 x 28 = 7843,136 (4 tokens)6,422,528 (8,192 tokens)23,520,000 (30,000 tokens)

Resizing behavior:

  • If the image pixel count is below min_pixels, the image is upscaled until it exceeds min_pixels.

  • If the image pixel count is within [min_pixels, max_pixels], the original image is used without resizing.

  • If the image pixel count exceeds max_pixels, the image is downscaled below max_pixels.

Output token limits

ModelDefault and maximummax_tokens
ModelDefault and maximummax_tokens
qwen-vl-ocr-latest, qwen-vl-ocr-2025-11-20, qwen-vl-ocr-2024-10-28Same as the model's maximum output length. See Supported models.
qwen-vl-ocr, qwen-vl-ocr-2025-04-13, qwen-vl-ocr-2025-08-284,096

For qwen-vl-ocr, qwen-vl-ocr-2025-04-13, and qwen-vl-ocr-2025-08-28, max_tokens defaults to 4096. To increase it (4097–8192), contact your commercial manager with: your Alibaba Cloud account ID, image type (e.g., documents, e-commerce, contracts), model name, estimated QPS and daily request volume, and the percentage of requests exceeding 4096 output tokens.


DashScope API

Endpoints

RegionHTTP endpoint
RegionHTTP endpoint
SingaporePOST https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
US (Virginia)POST https://dashscope-us.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
China (Beijing)POST https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation

SDK base URL configuration:

Python:

python
dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

Java (Method 1 -- constructor):

java
import com.alibaba.dashscope.protocol.Protocol;
MultiModalConversation conv = new MultiModalConversation(Protocol.HTTP.getValue(), "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1");

Java (Method 2 -- static block):

java
import com.alibaba.dashscope.utils.Constants;
Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";

Replace the domain with dashscope-us.aliyuncs.com for the US (Virginia) region or dashscope.aliyuncs.com for the China (Beijing) region. For the China (Beijing) region, you do not need to set base_url for SDK calls.

Get an API key and set it as an environment variable. If you use the DashScope SDK, you must also install the DashScope SDK.

Built-in tasks

The DashScope API provides built-in OCR tasks via the ocr_options parameter. Each task uses an optimized default prompt, eliminating the need for a text message.

Taskocr_options.taskvalueOutput format
Taskocr_options.taskvalueOutput format
General text recognitiontext_recognitionPlain text
High-precision recognitionadvanced_recognitionPlain text with bounding boxes
Information extractionkey_information_extractionStructured key-value pairs
Table parsingtable_parsingTable structure
Document parsingdocument_parsingDocument structure
Formula recognitionformula_recognitionLaTeX formulas
Multilingual recognitionmulti_lanMultilingual text

High-precision recognition

Returns text with positional data for each recognized line.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

messages = [{\
            "role": "user",\
            "content": [{\
                "image": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241108/ctdzex/biaozhun.jpg",\
                "min_pixels": 32 * 32 * 3,\
                "max_pixels": 32 * 32 * 8192,\
                "enable_rotate": False}]\
            }]

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model='qwen-vl-ocr-2025-11-20',
    messages=messages,
    ocr_options={"task": "advanced_recognition"}
)
print(response["output"]["choices"][0]["message"].content[0]["text"])

Java

java
// dashscope SDK version >= 2.21.8
import java.util.Arrays;
import java.util.Collections;
import java.util.Map;
import java.util.HashMap;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.aigc.multimodalconversation.OcrOptions;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.Constants;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241108/ctdzex/biaozhun.jpg");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        map.put("enable_rotate", false);

        OcrOptions ocrOptions = OcrOptions.builder()
                .task(OcrOptions.Task.ADVANCED_RECOGNITION)
                .build();
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map
                        )).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .ocrOptions(ocrOptions)
                .build();
        MultiModalConversationResult result = conv.call(param);
        System.out.println(result.getOutput().getChoices().get(0).getMessage().getContent().get(0).get("text"));
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '
{
  "model": "qwen-vl-ocr-2025-11-20",
  "input": {
    "messages": [\
      {\
        "role": "user",\
        "content": [\
          {\
            "image": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241108/ctdzex/biaozhun.jpg",\
            "min_pixels": 3072,\
            "max_pixels": 8388608,\
            "enable_rotate": false\
          }\
        ]\
      }\
    ]
  },
  "parameters": {
    "ocr_options": {
      "task": "advanced_recognition"
    }
  }
}
'

Information extraction

Extracts structured key-value data from images. Specify fields to extract in task_config.result_schema.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

messages = [\
      {\
        "role":"user",\
        "content":[\
          {\
              "image":"http://duguang-labelling.oss-cn-shanghai.aliyuncs.com/demo_ocr/receipt_zh_demo.jpg",\
              "min_pixels": 3072,\
              "max_pixels": 8388608,\
              "enable_rotate": False\
          }\
        ]\
      }\
    ]

params = {
  "ocr_options":{
    "task": "key_information_extraction",
    "task_config": {
      "result_schema": {
          "Ride Date": "Corresponds to the ride date and time in the image, in the format YYYY-MM-DD, for example, 2025-03-05",
          "Invoice Code": "Extract the invoice code from the image, usually a combination of numbers or letters",
          "Invoice Number": "Extract the number from the invoice, usually composed of only digits."
      }
    }
  }
}

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model='qwen-vl-ocr-2025-11-20',
    messages=messages,
    **params)

print(response.output.choices[0].message.content[0]["ocr_result"])

Java

java
import java.util.Arrays;
import java.util.Collections;
import java.util.Map;
import java.util.HashMap;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.aigc.multimodalconversation.OcrOptions;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.Constants;
import com.google.gson.JsonObject;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "http://duguang-labelling.oss-cn-shanghai.aliyuncs.com/demo_ocr/receipt_zh_demo.jpg");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        map.put("enable_rotate", false);

        JsonObject resultSchema = new JsonObject();
        resultSchema.addProperty("Ride Date", "Corresponds to the ride date and time in the image, in the format YYYY-MM-DD, for example, 2025-03-05");
        resultSchema.addProperty("Invoice Code", "Extract the invoice code from the image, usually a combination of numbers or letters");
        resultSchema.addProperty("Invoice Number", "Extract the number from the invoice, usually composed of only digits.");

        OcrOptions ocrOptions = OcrOptions.builder()
                .task(OcrOptions.Task.KEY_INFORMATION_EXTRACTION)
                .taskConfig(OcrOptions.TaskConfig.builder().resultSchema(resultSchema).build())
                .build();
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map
                        )).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .ocrOptions(ocrOptions)
                .build();
        MultiModalConversationResult result = conv.call(param);
        System.out.println(result.getOutput().getChoices().get(0).getMessage().getContent().get(0).get("ocr_result"));
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '
{
  "model": "qwen-vl-ocr-2025-11-20",
  "input": {
    "messages": [\
      {\
        "role": "user",\
        "content": [\
          {\
            "image": "http://duguang-labelling.oss-cn-shanghai.aliyuncs.com/demo_ocr/receipt_zh_demo.jpg",\
            "min_pixels": 3072,\
            "max_pixels": 8388608,\
            "enable_rotate": false\
          }\
        ]\
      }\
    ]
  },
  "parameters": {
    "ocr_options": {
      "task": "key_information_extraction",
      "task_config": {
        "result_schema": {
          "Ride Date": "Corresponds to the ride date and time in the image, in the format YYYY-MM-DD, for example, 2025-03-05",
          "Invoice Code": "Extract the invoice code from the image, usually a combination of numbers or letters",
          "Invoice Number": "Extract the number from the invoice, usually composed of only digits."
        }
      }
    }
  }
}
'

Table parsing

Extracts table structure from images.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

messages = [{\
            "role": "user",\
            "content": [{\
                "image": "http://duguang-llm.oss-cn-hangzhou.aliyuncs.com/llm_data_keeper/data/doc_parsing/tables/photo/eng/17.jpg",\
                "min_pixels": 32 * 32 * 3,\
                "max_pixels": 32 * 32 * 8192,\
                "enable_rotate": False}]\
            }]

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model='qwen-vl-ocr-2025-11-20',
    messages=messages,
    ocr_options={"task": "table_parsing"}
)
print(response["output"]["choices"][0]["message"].content[0]["text"])

Java

java
import java.util.Arrays;
import java.util.Collections;
import java.util.Map;
import java.util.HashMap;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.aigc.multimodalconversation.OcrOptions;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.Constants;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "https://duguang-llm.oss-cn-hangzhou.aliyuncs.com/llm_data_keeper/data/doc_parsing/tables/photo/eng/17.jpg");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        map.put("enable_rotate", false);

        OcrOptions ocrOptions = OcrOptions.builder()
                .task(OcrOptions.Task.TABLE_PARSING)
                .build();
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map
                        )).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .ocrOptions(ocrOptions)
                .build();
        MultiModalConversationResult result = conv.call(param);
        System.out.println(result.getOutput().getChoices().get(0).getMessage().getContent().get(0).get("text"));
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '
{
  "model": "qwen-vl-ocr-2025-11-20",
  "input": {
    "messages": [\
      {\
        "role": "user",\
        "content": [\
          {\
            "image": "http://duguang-llm.oss-cn-hangzhou.aliyuncs.com/llm_data_keeper/data/doc_parsing/tables/photo/eng/17.jpg",\
            "min_pixels": 3072,\
            "max_pixels": 8388608,\
            "enable_rotate": false\
          }\
        ]\
      }\
    ]
  },
  "parameters": {
    "ocr_options": {
      "task": "table_parsing"
    }
  }
}
'

Document parsing

Extracts the structural layout and text from documents.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

messages = [{\
            "role": "user",\
            "content": [{\
                "image": "https://img.alicdn.com/imgextra/i1/O1CN01ukECva1cisjyK6ZDK_!!6000000003635-0-tps-1500-1734.jpg",\
                "min_pixels": 32 * 32 * 3,\
                "max_pixels": 32 * 32 * 8192,\
                "enable_rotate": False}]\
            }]

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model='qwen-vl-ocr-2025-11-20',
    messages=messages,
    ocr_options={"task": "document_parsing"}
)
print(response["output"]["choices"][0]["message"].content[0]["text"])

Java

java
import java.util.Arrays;
import java.util.Collections;
import java.util.Map;
import java.util.HashMap;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.aigc.multimodalconversation.OcrOptions;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.Constants;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "https://img.alicdn.com/imgextra/i1/O1CN01ukECva1cisjyK6ZDK_!!6000000003635-0-tps-1500-1734.jpg");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        map.put("enable_rotate", false);

        OcrOptions ocrOptions = OcrOptions.builder()
                .task(OcrOptions.Task.DOCUMENT_PARSING)
                .build();
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map
                        )).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .ocrOptions(ocrOptions)
                .build();
        MultiModalConversationResult result = conv.call(param);
        System.out.println(result.getOutput().getChoices().get(0).getMessage().getContent().get(0).get("text"));
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '
{
  "model": "qwen-vl-ocr-2025-11-20",
  "input": {
    "messages": [\
      {\
        "role": "user",\
        "content": [\
          {\
            "image": "https://img.alicdn.com/imgextra/i1/O1CN01ukECva1cisjyK6ZDK_!!6000000003635-0-tps-1500-1734.jpg",\
            "min_pixels": 3072,\
            "max_pixels": 8388608,\
            "enable_rotate": false\
          }\
        ]\
      }\
    ]
  },
  "parameters": {
    "ocr_options": {
      "task": "document_parsing"
    }
  }
}
'

Formula recognition

Extracts mathematical formulas from images and returns them in LaTeX format.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

messages = [{\
            "role": "user",\
            "content": [{\
                "image": "http://duguang-llm.oss-cn-hangzhou.aliyuncs.com/llm_data_keeper/data/formula_handwriting/test/inline_5_4.jpg",\
                "min_pixels": 32 * 32 * 3,\
                "max_pixels": 32 * 32 * 8192,\
                "enable_rotate": False}]\
            }]

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model='qwen-vl-ocr-2025-11-20',
    messages=messages,
    ocr_options={"task": "formula_recognition"}
)
print(response["output"]["choices"][0]["message"].content[0]["text"])

Java

java
import java.util.Arrays;
import java.util.Collections;
import java.util.Map;
import java.util.HashMap;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.aigc.multimodalconversation.OcrOptions;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.Constants;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "http://duguang-llm.oss-cn-hangzhou.aliyuncs.com/llm_data_keeper/data/formula_handwriting/test/inline_5_4.jpg");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        map.put("enable_rotate", false);

        OcrOptions ocrOptions = OcrOptions.builder()
                .task(OcrOptions.Task.FORMULA_RECOGNITION)
                .build();
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map
                        )).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .ocrOptions(ocrOptions)
                .build();
        MultiModalConversationResult result = conv.call(param);
        System.out.println(result.getOutput().getChoices().get(0).getMessage().getContent().get(0).get("text"));
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '
{
  "model": "qwen-vl-ocr",
  "input": {
    "messages": [\
      {\
        "role": "user",\
        "content": [\
          {\
            "image": "http://duguang-llm.oss-cn-hangzhou.aliyuncs.com/llm_data_keeper/data/formula_handwriting/test/inline_5_4.jpg",\
            "min_pixels": 3072,\
            "max_pixels": 8388608,\
            "enable_rotate": false\
          }\
        ]\
      }\
    ]
  },
  "parameters": {
    "ocr_options": {
      "task": "formula_recognition"
    }
  }
}
'

General text recognition

Extracts plain text from images without structural formatting.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

messages = [{\
            "role": "user",\
            "content": [{\
                "image": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241108/ctdzex/biaozhun.jpg",\
                "min_pixels": 32 * 32 * 3,\
                "max_pixels": 32 * 32 * 8192,\
                "enable_rotate": False}]\
            }]

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model='qwen-vl-ocr-2025-11-20',
    messages=messages,
    ocr_options={"task": "text_recognition"}
)
print(response["output"]["choices"][0]["message"].content[0]["text"])

Java

java
import java.util.Arrays;
import java.util.Collections;
import java.util.Map;
import java.util.HashMap;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.aigc.multimodalconversation.OcrOptions;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.Constants;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241108/ctdzex/biaozhun.jpg");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        map.put("enable_rotate", false);

        OcrOptions ocrOptions = OcrOptions.builder()
                .task(OcrOptions.Task.TEXT_RECOGNITION)
                .build();
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map
                        )).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .ocrOptions(ocrOptions)
                .build();
        MultiModalConversationResult result = conv.call(param);
        System.out.println(result.getOutput().getChoices().get(0).getMessage().getContent().get(0).get("text"));
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '
{
  "model": "qwen-vl-ocr-2025-11-20",
  "input": {
    "messages": [\
      {\
        "role": "user",\
        "content": [\
          {\
            "image": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241108/ctdzex/biaozhun.jpg",\
            "min_pixels": 3072,\
            "max_pixels": 8388608,\
            "enable_rotate": false\
          }\
        ]\
      }\
    ]
  },
  "parameters": {
    "ocr_options": {
      "task": "text_recognition"
    }
  }
}
'

Multilingual recognition

Recognizes text in multiple languages from images.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

messages = [{\
            "role": "user",\
            "content": [{\
                "image": "https://img.alicdn.com/imgextra/i2/O1CN01VvUMNP1yq8YvkSDFY_!!6000000006629-2-tps-6000-3000.png",\
                "min_pixels": 32 * 32 * 3,\
                "max_pixels": 32 * 32 * 8192,\
                "enable_rotate": False}]\
            }]

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    model='qwen-vl-ocr-2025-11-20',
    messages=messages,
    ocr_options={"task": "multi_lan"}
)
print(response["output"]["choices"][0]["message"].content[0]["text"])

Java

java
import java.util.Arrays;
import java.util.Collections;
import java.util.Map;
import java.util.HashMap;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.aigc.multimodalconversation.OcrOptions;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.Constants;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "https://img.alicdn.com/imgextra/i2/O1CN01VvUMNP1yq8YvkSDFY_!!6000000006629-2-tps-6000-3000.png");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        map.put("enable_rotate", false);

        OcrOptions ocrOptions = OcrOptions.builder()
                .task(OcrOptions.Task.MULTI_LAN)
                .build();
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map
                        )).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .ocrOptions(ocrOptions)
                .build();
        MultiModalConversationResult result = conv.call(param);
        System.out.println(result.getOutput().getChoices().get(0).getMessage().getContent().get(0).get("text"));
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--data '
{
  "model": "qwen-vl-ocr-2025-11-20",
  "input": {
    "messages": [\
      {\
        "role": "user",\
        "content": [\
          {\
            "image": "https://img.alicdn.com/imgextra/i2/O1CN01VvUMNP1yq8YvkSDFY_!!6000000006629-2-tps-6000-3000.png",\
            "min_pixels": 3072,\
            "max_pixels": 8388608,\
            "enable_rotate": false\
          }\
        ]\
      }\
    ]
  },
  "parameters": {
    "ocr_options": {
      "task": "multi_lan"
    }
  }
}
'

Streaming (DashScope)

Enable streaming output to receive results incrementally. The method varies by SDK:

  • Python SDK: Set stream=True and incremental_output=True.

  • Java SDK: Use the streamCall interface.

  • HTTP: Set the X-DashScope-SSE: enable header.

Python

python
import os
import dashscope

dashscope.base_http_api_url = 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1'

PROMPT_TICKET_EXTRACTION = """
Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image.
You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?).
Return the data in JSON format as follows: {'invoice_number': 'xxx','departure_station': 'xxx', 'arrival_station': 'xxx', 'departure_date_and_time':'xxx', 'seat_number': 'xxx','ticket_price':'xxx', 'id_card_number': 'xxx', 'passenger_name': 'xxx'},
"""

messages = [\
    {\
        "role": "user",\
        "content": [\
            {\
                "image": "https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg",\
                "min_pixels": 32 * 32 * 3,\
                "max_pixels": 32 * 32 * 8192},\
            {\
                "type": "text",\
                "text": PROMPT_TICKET_EXTRACTION\
            }\
        ]\
    }\
]

response = dashscope.MultiModalConversation.call(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    model="qwen-vl-ocr-2025-11-20",
    messages=messages,
    stream=True,
    incremental_output=True,
)
full_content = ""
print("Streaming output content:")
for response in response:
    try:
        print(response["output"]["choices"][0]["message"].content[0]["text"])
        full_content += response["output"]["choices"][0]["message"].content[0]["text"]
    except:
        pass
print(f"Full content: {full_content}")

Java

java
import java.util.*;

import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.common.MultiModalMessage;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import io.reactivex.Flowable;
import com.alibaba.dashscope.utils.Constants;

public class Main {

    static {
        Constants.baseHttpApiUrl="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1";
    }

    public static void simpleMultiModalConversationCall()
            throws ApiException, NoApiKeyException, UploadFileException {
        MultiModalConversation conv = new MultiModalConversation();
        Map<String, Object> map = new HashMap<>();
        map.put("image", "https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg");
        map.put("max_pixels", 8388608);
        map.put("min_pixels", 3072);
        MultiModalMessage userMessage = MultiModalMessage.builder().role(Role.USER.getValue())
                .content(Arrays.asList(
                        map,
                        Collections.singletonMap("text", "Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image. You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?). Return the data in JSON format as follows: {\'invoice_number\': \'xxx\', \'departure_station\': \'xxx\', \'arrival_station\': \'xxx\', \'departure_date_and_time\':\'xxx\', \'seat_number\': \'xxx\',\'ticket_price\':\'xxx\', \'id_card_number\': \'xxx\', \'passenger_name\': \'xxx\'"))).build();
        MultiModalConversationParam param = MultiModalConversationParam.builder()
                .apiKey(System.getenv("DASHSCOPE_API_KEY"))
                .model("qwen-vl-ocr-2025-11-20")
                .message(userMessage)
                .incrementalOutput(true)
                .build();
        Flowable<MultiModalConversationResult> result = conv.streamCall(param);
        result.blockingForEach(item -> {
            try {
                List<Map<String, Object>> contentList = item.getOutput().getChoices().get(0).getMessage().getContent();
                if (!contentList.isEmpty()){
                    System.out.println(contentList.get(0).get("text"));
                }//
            } catch (Exception e){
                System.exit(0);
            }
        });
    }

    public static void main(String[] args) {
        try {
            simpleMultiModalConversationCall();
        } catch (ApiException | NoApiKeyException | UploadFileException e) {
            System.out.println(e.getMessage());
        }
        System.exit(0);
    }
}

curl

bash
curl --location 'https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \
--header "Authorization: Bearer $DASHSCOPE_API_KEY" \
--header 'Content-Type: application/json' \
--header 'X-DashScope-SSE: enable' \
--data '
{
    "model": "qwen-vl-ocr-2025-11-20",
    "input": {
        "messages": [\
            {\
              "role": "user",\
              "content": [\
                  {\
                      "image": "https://img.alicdn.com/imgextra/i2/O1CN01ktT8451iQutqReELT_!!6000000004408-0-tps-689-487.jpg",\
                      "min_pixels": 3072,\
                      "max_pixels": 8388608\
                  },\
                  {"type": "text", "text": "Please extract the invoice number, train number, departure station, arrival station, departure date and time, seat number, seat class, ticket price, ID card number, and passenger name from the train ticket image. You must accurately extract the key information. Do not omit or fabricate information. Replace any single character that is blurry or obscured by strong light with an English question mark (?). Return the data in JSON format as follows: {\'invoice_number\': \'xxx\', \'departure_station\': \'xxx\', \'arrival_station\': \'xxx\', \'departure_date_and_time\':\'xxx\', \'seat_number\': \'xxx\',\'ticket_price\':\'xxx\', \'id_card_number\': \'xxx\', \'passenger_name\': \'xxx\'}"}\
              ]\
            }\
        ]
    },
    "parameters": {
        "incremental_output": true
    }
}'

Request parameters

ParameterTypeRequiredDescription
ParameterTypeRequiredDescription
modelstringYesModel name. See Recommended models for supported models.
input.messagesarrayYesAn array of message objects.

Message object

Each message requires a role (must be user) and a content field (string or array). Use a string for text-only input. Use an array if the input includes image data, with these fields:

ParameterTypeRequiredDescription
ParameterTypeRequiredDescription
imagestringNoURL, Base64 Data URL, or local path of the image. See Passing local files.
textstringNoThe text prompt. Default: "Please output only the text content from the image without any additional descriptions or formatting". Not required when using a built-in task.
enable_rotatebooleanNoSet to true to correct skewed images. Default: false.
min_pixelsintegerNoMinimum pixel threshold. See Image resolution control.
max_pixelsintegerNoMaximum pixel threshold. See Image resolution control.

Generation parameters

Set these in the parameters object for HTTP calls.

ParameterTypeDefaultDescription
ParameterTypeDefaultDescription
max_tokensintegerVariesMaximum tokens in the output. See Output token limits. In the Java SDK, use maxTokens.
streambooleanfalseEnable streaming output. Python SDK only. For Java, use streamCall. For HTTP, set X-DashScope-SSE: enable.
incremental_outputbooleanfalseWhen true (recommended), each chunk contains only new content. When false, each chunk contains the full sequence so far. In the Java SDK, use incrementalOutput.
temperaturefloat0.01Controls output diversity. Range: [0, 2).
top_pfloat0.001Nucleus sampling threshold. Range: (0, 1.0]. Set either temperature or top_p, not both.
top_kinteger1Limits the candidate token set during sampling. If the value is None or greater than 100, the top_k policy is not enabled, and only the top_p policy takes effect. Must be >= 0.
repetition_penaltyfloat1.0Penalty for repeated sequences. Values above 1.0 reduce repetition.
presence_penaltyfloat0.0Controls content repetition. Range: [-2.0, 2.0].
seedinteger--Ensures reproducible results. Range: [0, 2^31 - 1].
logprobsbooleanfalseSet to true to return log probabilities. Supported models: qwen-vl-ocr-2025-04-13 and later. In the Java SDK, use the same name. For HTTP, place in parameters.
top_logprobsinteger0Number of most likely tokens per step. Range: [0, 5]. Only effective when logprobs is true. In the Java SDK, use topLogprobs. For HTTP, place in parameters.
stopstring or array--Stop words or token IDs. Generation stops when a specified string or token_id appears. Do not mix strings and token_ids in the same array.

Built-in task parameters (ocr_options )

When using a built-in task, pass ocr_options in parameters (HTTP), as a keyword argument (Python SDK), or via the OcrOptions builder (Java SDK).

ParameterTypeRequiredDescription
ParameterTypeRequiredDescription
ocr_options.taskstringYesBuilt-in task name. Valid values: text_recognition, key_information_extraction, document_parsing, table_parsing, formula_recognition, multi_lan, advanced_recognition.
ocr_options.task_configobjectNoConfiguration for key_information_extraction.
ocr_options.task_config.result_schemaobjectNoJSON object specifying fields to extract. Keys are field names, values are optional descriptions for improved accuracy. Supports up to three nesting levels.

result_schema example:

json
"result_schema": {
     "invoice_number": "The unique identification number of the invoice, usually a combination of numbers and letters.",
     "issue_date": "The date the invoice was issued. Extract it in YYYY-MM-DD format, for example, 2023-10-26.",
     "seller_name": "The full company name of the seller shown on the invoice.",
     "total_amount": "The total amount on the invoice, including tax. Extract the numerical value and keep two decimal places, for example, 123.45."
}

In the Java SDK, this parameter is OcrOptions. The minimum DashScope Python SDK version is 1.22.2. The minimum Java SDK version is 2.18.4. For advanced_recognition, Java SDK >= 2.21.8 is required.

Response

The DashScope API uses identical response format for streaming and non-streaming output.

json
{
  "status_code": 200,
  "request_id": "8f8c0f6e-6805-4056-bb65-d26d66080a41",
  "code": "",
  "message": "",
  "output": {
    "text": null,
    "finish_reason": null,
    "choices": [\
      {\
        "finish_reason": "stop",\
        "message": {\
          "role": "assistant",\
          "content": [\
            {\
              "ocr_result": {\
                "kv_result": {\
                  "price_excluding_tax": "230769.23",\
                  "invoice_code": "142011726001",\
                  "organization_code": "null",\
                  "buyer_name": "Cai Yingshi",\
                  "seller_name": "null"\
                }\
              },\
              "text": "```json\n{\n    \"price_excluding_tax\": \"230769.23\",\n    \"invoice_code\": \"142011726001\",\n    \"organization_code\": \"null\",\n    \"buyer_name\": \"Cai Yingshi\",\n    \"seller_name\": \"null\"\n}\n```"\
            }\
          ]\
        }\
      }\
    ],
    "audio": null
  },
  "usage": {
    "input_tokens": 926,
    "output_tokens": 72,
    "characters": 0,
    "image_tokens": 754,
    "input_tokens_details": {
      "image_tokens": 754,
      "text_tokens": 172
    },
    "output_tokens_details": {
      "text_tokens": 72
    },
    "total_tokens": 998
  }
}
FieldTypeDescription
FieldTypeDescription
status_codestring200 indicates success. The Java SDK throws an exception instead of returning this field.
request_idstringUnique request identifier. In the Java SDK, this is requestId.
codestringError code. Empty on success. Only the Python SDK returns this field.
output.textstringAlways null.
output.finish_reasonstringnull during generation, stop when complete, length when truncated.
output.choices[].finish_reasonstringSame values as output.finish_reason.
output.choices[].message.rolestringAlways assistant.
output.choices[].message.content[].textstringExtracted text or formatted output from the model.
output.choices[].message.content[].ocr_resultobjectReturned for built-in tasks (key_information_extraction, advanced_recognition).
output.choices[].message.content[].ocr_result.kv_resultobjectKey-value extraction results (for key_information_extraction).
output.choices[].message.content[].ocr_result.words_infoarrayText line results with positional data (for advanced_recognition).
output.choices[].message.content[].ocr_result.words_info[].rotate_rectarray[center_x, center_y, width, height, angle] -- rotated bounding rectangle.
output.choices[].message.content[].ocr_result.words_info[].locationarray[x1, y1, x2, y2, x3, y3, x4, y4] -- four vertices clockwise from the top-left.
output.choices[].message.content[].ocr_result.words_info[].textstringContent of the text line.
output.choices[].message.logprobsobjectLog probability information, returned when logprobs is true.
usage.input_tokensintegerInput token count.
usage.output_tokensintegerOutput token count.
usage.charactersintegerFixed to 0.
usage.total_tokensintegerSum of input_tokens and output_tokens.
usage.image_tokensintegerTokens corresponding to the image input.
usage.input_tokens_details.image_tokensintegerImage input tokens.
usage.input_tokens_details.text_tokensintegerText input tokens.
usage.output_tokens_details.text_tokensintegerText output tokens.

Supported models

ModelDescription
ModelDescription
qwen-vl-ocr-latestAlways points to the latest version.
qwen-vl-ocr-2025-11-20Latest dated snapshot.
qwen-vl-ocr-2025-08-28Previous version.
qwen-vl-ocr-2025-04-13Previous version.
qwen-vl-ocr-2024-10-28Previous version.
qwen-vl-ocrBase model.

Error codes

If a model call returns an error, see Error messages to resolve the issue.

Previous: Qwen-Deep-ResearchNext: OpenAI-compatible - Chat

Is this page helpful?

OpenAI-compatible API

Endpoints

Prerequisites

Quick start

Request parameters

Response

Image resolution control

Output token limits

DashScope API

Endpoints

Built-in tasks

Streaming (DashScope)

Request parameters

Response

Supported models

Error codes

Contact Us

Sales Support

Live-chat with our sales team or get in touch with a business development professional in your region.

Contact Sales

Technical Support

Open a ticket and get quick help from our technical team.

Open a Ticket >

Connect & Report Abuse

We look forward to your suggestion.

Post a Suggestion > Report Abuse >

\ \ Hi, I'm Alibaba Cloud AI Assistant!\ \ I can help with questions and solutions.

Why Alibaba Cloud

About Alibaba Cloud

Asia Accelerator

Our Global Network

Global Offices

Trust Center

Case Studies

Analyst Reports

Products & Pricings

Pricing Calculator

ECS

SAS

Model Studio

Database

Security

SMS

Solutions

Financial Services

Retail Services

Media Services

Gaming Services

ISV Solutions

Engage

Developer Community

Partner Network

Startups

Marketplace

Join Alibaba Cloud

Resources & Support

Developer Learning Hub

Documentation Center

Training & Certification

Service Notices

Submit a Ticket

Security Report

Qwen Cloud

Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links

© 2009-2026 Copyright by Alibaba Cloud All rights reserved

© 2009-2026 Copyright by Alibaba Cloud All rights reserved

Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links