Skip to content

Context cache for Qwen models

Reference, synced 2026-06-13.

flowchart TD
  n0["Models"]
  n1["Overview"]
  n2["Products"]
  n3["Solutions"]
  n4["Pricing"]
  n5["Resources"]
  n6["Partners"]
  n7["Support"]
  n8["Language"]
  n0 --> n1
  n1 --> n2
  n2 --> n3
  n3 --> n4
  n4 --> n5
  n5 --> n6
  n6 --> n7
  n7 --> n8

Free access Accelerate Delivery with Fixed-Cost Agentic CodingWatch how it works

Models

Empowering AI innovation for both enterprises and developers with Alibaba Cloud’s best-in-class Qwen models, AI-native apps, and AI solutions.

Alibaba Cloud Model Studio \ Enterprise-grade large model service and application development platform.

Try Visual Model \ Supports image understanding, image generation, and video generation.

Models

HappyHorse-1.0-T2V \ Cinematic creative generation, ultimate dynamic details Qwen3-VL-Plus \ Native VL, spatial reasoning, 1M-context video analysis Wan2.7-VideoEdit \ Supports both localized and global editing with prompt

Qwen3.6-Plus \ Native multimodal, 1M context, agentic coding Wan2.7-Image-Pro \ Interactive editing, long-text rendering, precise prompt following Qwen-Plus \ Balanced intelligence, efficient inference, production-ready performance

Qwen-Image-2.0 \ Professional infographics, exquisite photorealism Z-Image-Turbo \ Ultra-fast image generation, high throughput, cost-optimized inference Qwen3-Coder-Next \ Multi-turn tool interactions, future-ready development support

Wan2.7-T2V \ High-fidelity T2V, 15s duration, advanced camera control Wan2.7-I2V \ Cinematic I2V with emotional depth and visceral impact Wan2.7-R2V \ Up to 5 mixed image/video inputs and audio timbre cloning

GenAI Application

Qoder \ Intelligent coding assistant, available for enterprise-dedicated deployment. Qoder CN \ AI-powered coding assistant that boosts developer productivity with intelligent code completion, AI chat, multi-file editing, and task automation.

AI Service

Model Experience \ Experience full-scale, multimodal model capabilities online. Platform for AI \ An AI-native algorithm engineering platform for end-to-end modeling, training, and inference service deployment. Fine-tune Video Generation Model \ Customize Wan’s text-to-video capabilities through model fine-tuning to meet your unique requirements.

AI Use Case

AI Savings Plan Hot \ Save up to 47% on AI costs. Limited-time offer tailored to your usage. AI Video Creation \ Elevate your professional video production with Wan 2.6.

AI Token Plan \ One plan. Multiple models. Big Savings with a Fixed Subscription. AI Image Creation \ All-in-one creative suite for copywriting, image generation, and poster design.

Overview

As a global full-stack AI leader, Alibaba Cloud aims to make computing accessible to everyone and help worldwide customers accelerate innovation.

Why Alibaba Cloud

About Alibaba Cloud \ AI Powered Cloud Technology Our Global Network \ Explore our global presence and deployment regions around the world Our Global Offices \ With offices in 4 continents, we're always close to where it matters.

Asia Accelerator \ Accelerate Success in Asia with Alibaba Cloud Go Global \ Benefits of our Global Alliance Trust Center \ Empowering enterprises with a secure, compliant, and globally trusted cloud infrastructure

Customers and Insights

Olympic Games \ Alibaba Cloud Powers Olympic Games with AI-powered cloud technology Case Studies \ Learn how customers are scaling their businesses on Alibaba Cloud Analyst Reports \ Learn what the top industry analyst firms are saying about Alibaba Cloud

What's New

Events and Webinars \ Quick access to upcoming and on-demand events Product Updates \ Stay informed of the latest innovations Press Room \ Latest news and media releases

Products

Featured ProductsAI & Machine Learning Computing Container Storage Networking & CDN Security Middleware Database Analytics ComputingMedia ServicesEnterprise Services & Cloud CommunicationDomain Names and WebsitesEnd User ComputingServerlessDeveloper ToolsMigration & O&M ManagementApsara Stack

Alibaba Cloud Model Studio \ Supercharge your AI journey effortlessly with industry-leading GenAI models ApsaraDB RDS \ Store and manage your business data, with automated monitoring and backups Certificate Management Service (Original SSL Certificate) \ Create a safe and secure connection between your website and users

Elastic Compute Service (ECS) \ Host websites anywhere and scale enterprise workloads Container Service for Kubernetes (ACK) \ Run and scale containerized applications on managed Kubernetes infrastructure Object Storage Service (OSS) \ Store large amounts of data in the cloud and access it anywhere, anytime

Simple Application Server (SAS) \ All-in-one services for fast deployment Elastic IP Address (EIP) \ Manage your public IPs independently to improve internet network quality Domain Names and Website \ Get the perfect domain name to suit your every need

Solutions

Solutions by Industry Technical Solutions AI WebsitesNetworking Security and ComplianceData and AnalyticsEnterprise Service and ApplicationCloud MigrationCloud NativeHybrid CloudSMB solutions

Financial Services \ Innovate faster with Alibaba Cloud Games \ Grow your game rapidly with high global availability

New Retail \ Alibaba Cloud enables digital retail transformation to fuel growth and realize an omnichannel customer experience throughout the consumer journey. Media and Entertainment \ Ready your content for today's media market with a digitalized media journey

Supply Chain \ Power your supply chain with intelligent, efficient, and reliable solutions Sports \ Digitizing the sports industry with intelligent tech

Sustainability \ Achieve a sustainable future with low-carbon and energy-efficient technologies

Pricing

Flexible options like pay-as-you-go and clear billing rules to meet diverse business needs.

Overview & Tools

Pricing Calculator \ Get an instant pricing estimate based on your usage and needs Free Trial \ Try our 80+ cloud products for free.

Pricing Options \ Get the most out of Alibaba Cloud with flexible pricing

Optimize your cost

Migrate & Save \ Superior Performance At Lower Pricing. Save up to 50%. Promotion Center \ Unlock the latest Alibaba Cloud offers & promos

Resources

Official documentation, extensive tools, training resources, and a community to grow and innovate in the cloud.

Technical Resources

Documentation \ Product guides and FAQs Architecture Center \ Design reliable, secure, and efficient cloud architecture. Intelligent Solution Explorer \ Find the right solution for you, powered by AI

Blog \ Latest cloud insights and developer trends Whitepapers \ Research that explores the how and why behind our technology

Training&Certification

Alibaba Cloud Academy \ Build cloud skills and earn certifications with expert-led training.

Developer Hub

Alibaba Cloud Project Hub \ Explore real-world projects built by developers using our platform. Our Developer MVPs \ Celebrating the developers who lead, build, and inspire our community

Partners

Partner-first strategy offering collaborative product, sales, and service models, plus high-quality partner solutions that complement Alibaba Cloud’s capabilities.

Marketplace

AI Alliance for ISVs \ Partner with us to build and grow AI solutions together ISV Benefits \ Unlock resources, market access, and go-to-market support as an ISV partner

Alibaba Cloud Marketplace \ Explore ready-to-deploy solutions from our partners and ISVs

Find a Partner

Partner Hub \ Find your ideal partner in no time

Become a Partner

Partner Network \ A partner portal for Alibaba Cloud Channel, Technology, MSP partner and other partner programs

Support

Full-lifecycle support and expert services, from cloud advisory and migration to operations.

Support & Professional Services

Professional Services \ Expert-led services to design, migrate, and optimize your cloud journey Support Plans \ Flexible support for every stage — from startup to enterprise

Partner Support Program \ Priority technical support for partners, with dedicated managers and faster issue resolution

Contact us

Connect With Us \

Talk to a sales expert and get a custom quote for your business

Language

  • English
  • 简体中文
  • 繁體中文
  • 日本語
  • Bahasa Indonesia

Locale

Visit aliyun.com

Documentation

Alibaba Cloud Model Studio

User Guide (Models) User Guide (Application) API Reference (Models) API Reference (Application)

Search for Help Content

Getting Started

The Beginner's Guide

Well-Architected Framework

AI & Machine Learning

Platform For AI

Alibaba Cloud Model Studio

DashVector

Artificial Intelligence Recommendation

OpenSearch

Image Search

Machine Translation

Intelligent Speech Interaction

Optimization Solver

Intelligent Computing LINGJUN

Computing

Elastic Compute Service

Elastic GPU Service

Elastic Container Instance

Dedicated Host

Compute Nest

Simple Application Server

Cloud Box

Auto Scaling

Elastic High Performance Computing

Batch Compute (Deprecated)

Function Compute

Serverless App Engine

ENS

Elastic Desktop Service

App Streaming

WUYING Terminal

Cloud Phone

Edge Network Acceleration

Alibaba Cloud Linux

AgentBay

Container

Container Service for Kubernetes

Container Compute Service

Container Registry

Storage

Object Storage Service

Cloud Parallel File Storage

File Storage NAS

Tablestore

Storage Capacity Unit

Simple Log Service

Cloud Backup

Intelligent Media Management

Drive and Photo Service

Data Transport

Cloud Storage Gateway

Data Online Migration

Hybrid Cloud Storage Array

Storage Services Overview

Backup and Disaster Recovery Center

Networking and CDN

Server Load Balancer

Elastic IP Address

Internet Shared Bandwidth

Data Transfer Plan

Virtual Private Cloud

NAT Gateway

PrivateLink

Alibaba Cloud DNS PrivateZone

Network Intelligence Service

Cloud Data Transfer

IPv6 Gateway

Anycast Elastic IP Address

Cloud Enterprise Network

Global Accelerator

VPN Gateway

Smart Access Gateway

Express Connect

CDN

Edge Security Acceleration

Cloud Network Well-architected Design Guidelines

Security

Anti-DDoS

Web Application Firewall

Cloud Firewall

Security Center

Bastionhost

Secure Access Service Edge

Certificate Management Service

Key Management Service

Data Security Center

Identity as a Service

Fraud Detection

AI Guardrails

Captcha

Blockchain as a Service

ID Verification

Managed Security Service

Middleware

Enterprise Distributed Application Service

Microservices Engine

Alibaba Cloud Service Mesh

SchedulerX

ApsaraMQ for RocketMQ

ApsaraMQ for Kafka

ApsaraMQ for RabbitMQ

ApsaraMQ for MQTT

Simple Message Queue (formerly MNS)

CloudFlow

EventBridge

Application Real-Time Monitoring Service

Managed Service for Prometheus

Managed Service for Grafana

Managed Service for OpenTelemetry

Performance Testing

STAROps

Databases

ApsaraDB Console

PolarDB

ApsaraDB RDS

ApsaraDB for OceanBase (Deprecated)

Tair (Redis® OSS-Compatible)

Lindorm

Time Series Database

ApsaraDB for MongoDB

ApsaraDB for HBase

ApsaraDB for Memcache

ApsaraDB for MyBase

AnalyticDB

ApsaraDB for ClickHouse

ApsaraDB for SelectDB

Data Transmission Service

Database Autonomy Service

Data Management

Database Gateway - Deprecated

ApsaraDB for Cassandra - Deprecated

Analytics Computing

MaxCompute

Hologres

Realtime Compute for Apache Flink

Elasticsearch

Vector Retrieval Service for Milvus

E-MapReduce

Data Lake Formation

DataV

Quick BI

Quick Audience

Quick Tracking

DataWorks

DataHub

Dataphin

Media Services

ApsaraVideo VOD

ApsaraVideo Live

Intelligent Media Services

ApsaraVideo Media Processing

Apsara Video SDK

Enterprise Services & Cloud Communication

Energy Expert

CloudQuotation

Salesforce on Alibaba Cloud

GoChina ICP Filing Assistant

Marketplace

Alibaba Mail

Direct Mail

Short Message Service

Voice Service

Phone Number Verification Service

Cell Phone Number Service

Chat App Message Service

Financial Intelligence Engine

Domain Names and Websites

Domain Names

ICP Filing

Alibaba Cloud DNS

End User Computing

Elastic Desktop Service

App Streaming

WUYING Terminal

Cloud Phone

AgentBay

Internet of Things

IoT Platform

Serverless

Serverless App Engine

CloudFlow

EventBridge

Simple Message Queue (formerly MNS)

Function Compute

Developer Tools

OpenAPI Explorer

Alibaba Cloud SDK

Cloud Shell

Resource Orchestration Service

Alibaba Cloud CLI

BSS OpenAPI

Terraform

Pulumi

Ticket System API

Mobile Platform as a Service

Alibaba Cloud DevOps

API Gateway

Cloud Control API

AI Coding Assistant Lingma

Cloud Skills Portal

Migration & O&M Management

CloudOps Orchestration Service

Cloud Monitor

Intelligent Advisor

Cloud Governance Center

ActionTrail

Cloud Config

Resource Access Management

Resource Management

Cloud Architect Design Tools

Migration Hub

Server Migration Center

Service Catalog

Logic Composer

Quota Center

CloudSSO

HTTPDNS

Solutions

SAP

SuperApp

OpenLake

Membership Service

Expenses and Costs

Account Center

More

Support

Legal

Tech Share Terms and Conditions

After Sales Support

China Gateway Program

Service Level Objectives

Management Console

Security Control

Inference requests to a large model often include overlapping input, such as in a multi-turn conversation or a series of questions about the same book. Context Cache caches the common prefix of these requests to reduce redundant computation during inference. This improves response speed and lowers usage costs without affecting response quality.

To accommodate various scenarios, the context cache offers two work modes. Choose a mode based on your requirements for convenience, determinism, and cost:

  • Explicit cache: A cache mode that you must actively enable. You create a cache for specific content to ensure a deterministic hit within its 5-minute validity period. Tokens used to create the cache are billed at 125% of the standard input token price, while subsequent cache hits are billed at 10% of that price.

  • Implicit cache: This automatic mode requires no configuration and cannot be disabled, making it ideal for scenarios that prioritize convenience. The system automatically identifies and caches the common prefix of requests, but cache hits are not guaranteed. The portion of the input served from the cache is billed at 20% of the standard input token price.

ItemExplicit cacheImplicit cache
ItemExplicit cacheImplicit cache
Affects response qualityNoNo
Billing for cache creation tokens125% of the standard input token price100% of the standard input token price
Billing for cached input tokens10% of the standard input token price20% of the standard input token price
Minimum tokens for caching1024256
Cache validity period5 minutes (resets on hit)Not guaranteed; the system periodically clears inactive data.

Note

Explicit cache and implicit cache are mutually exclusive. A request can use only one work mode.

Note

This topic covers OpenAI Chat Completions, DashScope, and Anthropic-compatible interfaces. Use the session cache with the Responses API to reduce inference latency and cost. For details, see Session cache.

Explicit cache

Compared to implicit cache, explicit cache requires manual creation and incurs an initial overhead. However, it provides a higher cache hit ratio and lower access latency.

How it works

To use explicit cache, add a "cache_control": {"type": "ephemeral"} marker in your messages array. The system then traces back from the position of each cache_control marker, examining up to 20 preceding content blocks to attempt a cache hit.

A single request supports up to four cache markers.

  • Cache miss

If a cache miss occurs, the system creates a new cache block from the content between the start of the messages array and the cache_control marker. The new cache block is valid for 5 minutes.

Cache creation occurs after the model generates a response. We recommend waiting for the creation request to complete before attempting to hit that cache.

A cache block contains at least 1024 tokens.

  • Cache hit

If a cache hit occurs, the system selects the longest matching prefix as the cache block and resets its validity period to 5 minutes.

The following example demonstrates how this works:

  1. Send the first request: Send a system message that contains text (A) with more than 1,024 tokens and add a cache marker.
json
[{"role": "system", "content": [{"type": "text", "text": A, "cache_control": {"type": "ephemeral"}}]}]

The system creates the first cache block, called cache block A.

  1. Send the second request: Send a request with the following structure.
json
[\
       {"role": "system", "content": A},\
       <Other messages>\
       {"role": "user","content": [{"type": "text", "text": B, "cache_control": {"type": "ephemeral"}}]}\
]
  • If there are 20 or fewer "Other messages," the request hits cache block A, and its validity period is reset to 5 minutes. The system also creates a new cache block based on A, the other messages, and B.

  • If there are more than 20 "Other messages," the request misses cache block A. The system still creates a new cache block based on the full context, which includes A, the other messages, and B.

Supported models

Singapore

China (Beijing)

Germany (Frankfurt)

Hong Kong (China)

The following models are all in the International deployment scope.

Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08, qwen3.6-max-preview, qwen3-max

Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen3.6-plus, qwen3.5-plus, qwen3.5-plus-2026-04-20, qwen-plus

Qwen Flash: qwen3.6-flash, qwen3.5-flash, qwen-flash

Qwen Coder: qwen3-coder-plus, qwen3-coder-flash

Qwen VL: qwen3-vl-plus, qwen3-vl-flash

DeepSeek: deepseek-v3.2

The following models are all in the Chinese mainland deployment scope.

Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08, qwen3.6-max-preview, qwen3-max

Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen3.6-plus, qwen3.5-plus, qwen3.5-plus-2026-04-20, qwen-plus

Qwen Flash: qwen3.6-flash, qwen3.5-flash, qwen-flash

Qwen Coder: qwen3-coder-plus, qwen3-coder-flash

Qwen VL: qwen3-vl-plus, qwen3-vl-flash

DeepSeek: deepseek-v3.2

Kimi: kimi-k2.6, kimi-k2.5

GLM: glm-5.1

Different deployment scopes support different models.

  • Global deployment scope:

Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08, qwen3-max

Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen3.6-plus, qwen3.5-plus, qwen-plus

Qwen Flash: qwen3.6-flash, qwen3.5-flash, qwen-flash

Qwen VL: qwen3-vl-plus

Qwen Coder: qwen3-coder-plus, qwen3-coder-flash

Kimi: kimi-k2.5

  • EU deployment scope:

Qwen Max: qwen3-max

Qwen Plus: qwen-plus

Qwen Flash: qwen3.6-flash, qwen3.5-flash

Qwen VL: qwen3-vl-plus

Different deployment scopes support different models.

  • Global deployment scope:

Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08

Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen3.6-plus

Qwen Flash: qwen3.6-flash

  • Hong Kong (China) deployment scope:

Qwen Max: qwen3-max

Qwen Plus: qwen-plus

Qwen Flash: qwen3.6-flash, qwen3.5-flash

Qwen VL: qwen3-vl-plus

Quick start

These examples demonstrate how cache blocks are created and hit using OpenAI-compatible, DashScope, and Anthropic-compatible protocols.

OpenAI compatible

DashScope

Anthropic compatible

python
from openai import OpenAI
import os

client = OpenAI(
    # If the environment variable is not set, replace the following line with: api_key="sk-xxx"
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    # If using a model in the China (Beijing) region, replace the base_url with: https://dashscope.aliyuncs.com/compatible-mode/v1
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)

# Mock code repository content. The minimum cacheable prompt length is 1,024 tokens.
long_text_content = "<Your Code Here>" * 400

# Function to send a request.
def get_completion(user_input):
    messages = [\
        {\
            "role": "system",\
            "content": [\
                {\
                    "type": "text",\
                    "text": long_text_content,\
                    # Place the cache_control marker here. This creates a cache block containing all content from the start of the messages array to the current content's position.\
                    "cache_control": {"type": "ephemeral"},\
                }\
            ],\
        },\
        # The user's question is different for each request.\
        {\
            "role": "user",\
            "content": user_input,\
        },\
    ]
    completion = client.chat.completions.create(
        # Select a model that supports explicit cache.
        model="qwen3-coder-plus",
        messages=messages,
    )
    return completion

# First request
first_completion = get_completion("What is the content of this code?")
print(f"First request cache creation tokens: {first_completion.usage.prompt_tokens_details.cache_creation_input_tokens}")
print(f"First request cached tokens: {first_completion.usage.prompt_tokens_details.cached_tokens}")
print("=" * 20)
# Second request. The code content is the same, only the question is different.
second_completion = get_completion("How can this code be optimized?")
print(f"Second request cache creation tokens: {second_completion.usage.prompt_tokens_details.cache_creation_input_tokens}")
print(f"Second request cached tokens: {second_completion.usage.prompt_tokens_details.cached_tokens}")

Python

Java

Python

python
import os
from dashscope import Generation
# If using a model in the China (Beijing) region, replace the base_url with: https://dashscope.aliyuncs.com/api/v1
dashscope.base_http_api_url = "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1"

# Mock code repository content. The minimum cacheable prompt length is 1,024 tokens.
long_text_content = "<Your Code Here>" * 400

# Function to send a request.
def get_completion(user_input):
    messages = [\
        {\
            "role": "system",\
            "content": [\
                {\
                    "type": "text",\
                    "text": long_text_content,\
                    # Place the cache_control marker here. This creates a cache block containing all content from the start of the messages array to the current content's position.\
                    "cache_control": {"type": "ephemeral"},\
                }\
            ],\
        },\
        # The user's question is different for each request.\
        {\
            "role": "user",\
            "content": user_input,\
        },\
    ]
    response = Generation.call(
        # If the environment variable is not set, replace the following line with your Model Studio API key: api_key = "sk-xxx",
        api_key=os.getenv("DASHSCOPE_API_KEY"),
        model="qwen3-coder-plus",
        messages=messages,
        result_format="message"
    )
    return response

# First request
first_completion = get_completion("What is the content of this code?")
print(f"First request cache creation tokens: {first_completion.usage.prompt_tokens_details['cache_creation_input_tokens']}")
print(f"First request cached tokens: {first_completion.usage.prompt_tokens_details['cached_tokens']}")
print("=" * 20)
# Second request. The code content is the same, only the question is different.
second_completion = get_completion("How can this code be optimized?")
print(f"Second request cache creation tokens: {second_completion.usage.prompt_tokens_details['cache_creation_input_tokens']}")
print(f"Second request cached tokens: {second_completion.usage.prompt_tokens_details['cached_tokens']}")

Java

java
// Minimum Java SDK version: 2.21.6
import com.alibaba.dashscope.aigc.generation.Generation;
import com.alibaba.dashscope.aigc.generation.GenerationParam;
import com.alibaba.dashscope.aigc.generation.GenerationResult;
import com.alibaba.dashscope.common.Message;
import com.alibaba.dashscope.common.MessageContentText;
import com.alibaba.dashscope.common.Role;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.InputRequiredException;
import com.alibaba.dashscope.exception.NoApiKeyException;

import java.util.Arrays;
import java.util.Collections;

public class Main {
    private static final String MODEL = "qwen3-coder-plus";
    // Mock code repository content (repeated 400 times to exceed 1,024 tokens).
    private static final String LONG_TEXT_CONTENT = generateLongText(400);
    private static String generateLongText(int repeatCount) {
        StringBuilder sb = new StringBuilder();
        for (int i = 0; i < repeatCount; i++) {
            sb.append("<Your Code Here>");
        }
        return sb.toString();
    }
    private static GenerationResult getCompletion(String userQuestion)
            throws NoApiKeyException, ApiException, InputRequiredException {
        // If using a model in the China (Beijing) region, replace the base_url with: https://dashscope.aliyuncs.com/api/v1
        Generation gen = new Generation("http", "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1");

        // Build the system message with cache control.
        MessageContentText systemContent = MessageContentText.builder()
                .type("text")
                .text(LONG_TEXT_CONTENT)
                .cacheControl(MessageContentText.CacheControl.builder()
                        .type("ephemeral") // Set the cache type.
                        .build())
                .build();

        Message systemMsg = Message.builder()
                .role(Role.SYSTEM.getValue())
                .contents(Collections.singletonList(systemContent))
                .build();
        Message userMsg = Message.builder()
                .role(Role.USER.getValue())
                .content(userQuestion)
                .build();

        // Build the request parameters.
        GenerationParam param = GenerationParam.builder()
                .model(MODEL)
                .messages(Arrays.asList(systemMsg, userMsg))
                .resultFormat(GenerationParam.ResultFormat.MESSAGE)
                .build();
        return gen.call(param);
    }

    private static void printCacheInfo(GenerationResult result, String requestLabel) {
        System.out.printf("%s cache creation tokens: %d%n", requestLabel, result.getUsage().getPromptTokensDetails().getCacheCreationInputTokens());
        System.out.printf("%s cached tokens: %d%n", requestLabel, result.getUsage().getPromptTokensDetails().getCachedTokens());
    }

    public static void main(String[] args) {
        try {
            // First request
            GenerationResult firstResult = getCompletion("What is the content of this code?");
            printCacheInfo(firstResult, "First request");
            System.out.println(new String(new char[20]).replace('\0', '='));
            // Second request
            GenerationResult secondResult = getCompletion("How can this code be optimized?");
            printCacheInfo(secondResult, "Second request");
        } catch (NoApiKeyException | ApiException | InputRequiredException e) {
            System.err.println("API call failed: " + e.getMessage());
            e.printStackTrace();
        }
    }
}
python
import anthropic
import os

client = anthropic.Anthropic(
    # If the environment variable is not set, replace the following line with: api_key="sk-xxx"
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    # If using a model in the China (Beijing) region, replace the base_url with: https://dashscope.aliyuncs.com/apps/anthropic
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/apps/anthropic",
)

# Mock code repository content. The minimum cacheable prompt length is 1,024 tokens.
long_text_content = "<Your Code Here>" * 400

# Function to send a request.
def get_completion(user_input):
    response = client.messages.create(
        # Select a model that supports explicit cache.
        model="qwen3-coder-plus",
        max_tokens=1024,
        system=[\
            {\
                "type": "text",\
                "text": long_text_content,\
                # Place the cache_control marker in 'text' content to create a cache. This can be in a system message (as shown) or in user/assistant/tool messages.\
                "cache_control": {"type": "ephemeral"},\
            }\
        ],
        messages=[\
            # The user's question is different for each request.\
            {"role": "user", "content": user_input},\
        ],
    )
    return response

# First request
first_completion = get_completion("What is the content of this code?")
print(f"First request cache creation tokens: {first_completion.usage.cache_creation_input_tokens}")
print(f"First request cached tokens: {first_completion.usage.cache_read_input_tokens}")
print("=" * 20)
# Second request. The code content is the same, only the question is different.
second_completion = get_completion("How can this code be optimized?")
print(f"Second request cache creation tokens: {second_completion.usage.cache_creation_input_tokens}")
print(f"Second request cached tokens: {second_completion.usage.cache_read_input_tokens}")

Adding the cache_control marker to the mock code repository content enables explicit cache. For subsequent requests about this repository, the system can reuse the cache block to avoid recomputation. This results in faster responses and lower costs compared to requests that do not hit the cache.

plaintext
First request cache creation tokens: 1605
First request cached tokens: 0
====================
Second request cache creation tokens: 0
Second request cached tokens: 1605

Fine-grained cache control

In complex scenarios, prompts often consist of multiple parts with different reuse frequencies. You can use multiple cache markers to achieve fine-grained control.

For example, the prompt for a smart customer service agent typically includes:

  • System persona: Highly stable and rarely changes.

  • External knowledge: Semi-stable. It is obtained through knowledge base retrieval or tool queries and might remain unchanged within a single conversation.

  • Conversation history: Grows dynamically.

  • Current question: Different for each request.

If you cache the entire prompt as a single unit, any minor change, such as an update to the external knowledge, can cause a cache miss.

You can add up to four cache markers in a request to create separate cache blocks for different parts of the prompt. This improves the cache hit ratio and enables fine-grained control.

Billing

Explicit cache only affects the billing of input tokens. The rules are as follows:

  • Cache creation: Content used to create a new cache is billed at 125% of the standard input token price. If the content for a new cache includes an existing cache as a prefix, only the additional part is billed for cache creation (i.e., new cache tokens minus existing cache tokens).

For example, if you have an existing 1,200-token cache (Cache A) and a new request needs to cache 1,500 tokens of content (Content AB), the first 1,200 tokens are billed as a cache hit at 10% of the standard price. The new 300 tokens are billed for cache creation at 125% of the standard price.

You can view the number of tokens used for cache creation in the cache_creation_input_tokens parameter.

  • Cache hit: Billed at 10% of the standard input token price.

You can view the number of cached tokens in the cached_tokens parameter (or cache_read_input_tokens for the Anthropic-compatible protocol).

  • Other tokens: Tokens that are neither a cache hit nor used for cache creation are billed at the standard input token price.

Cacheable content

Only the following message types in the messages array support cache markers:

  • system message

Note

If a request includes the tools parameter for a function calling scenario, the tool definition is included as part of the system message for cache calculation. Tool definitions cannot be cached independently. Cache markers added to tool definitions are ignored, as they can only be added to the content of messages.

  • user message

When you create a cache with the qwen3-vl-plus model, the cache_control marker can be placed after multimodal content or text. Its position does not affect how the entire user message is cached.

  • assistant message

  • tool message (the result after tool execution)

For a system message, for example, you must change the content field to an array and add the cache_control field:

json
{
  "role": "system",
  "content": [\
    {\
      "type": "text",\
      "text": "<your specified prompt>",\
      "cache_control": {\
        "type": "ephemeral"\
      }\
    }\
  ]
}

This structure also applies to other message types in the messages array.

Limitations

  • The minimum cacheable prompt length is 1,024 tokens.

  • The cache uses a prefix matching strategy, searching backward from a cache_control marker through up to 20 preceding content blocks. A cache hit cannot occur if the match is outside this range.

  • The type parameter must be set to ephemeral, which has a validity period of 5 minutes.

  • A single request supports up to four cache markers.

If more than four cache markers are provided, only the last four take effect.

Function calling cache optimization

Because a tool definition is serialized into a JSON string for cache calculation, you must ensure that the tool definition is identical across requests to avoid cache invalidation. Pay close attention to the following:

  • Consistent tool list order: The order of tools in the tools array must be consistent.

  • Consistent field order: The order of JSON fields for the same tool must be consistent.

  • Consistent field structure: Do not omit or add fields, even if a field is empty or optional.

Usage examples

Asking different questions about a long text

python
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    # If using a model in the China (Beijing) region, replace the base_url with: https://dashscope.aliyuncs.com/compatible-mode/v1
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)

# Mock code repository content.
long_text_content = "<Your Code Here>" * 400

# Function to send a request.
def get_completion(user_input):
    messages = [\
        {\
            "role": "system",\
            "content": [\
                {\
                    "type": "text",\
                    "text": long_text_content,\
                    # Place the cache_control marker here to create a cache from the start of the prompt to this content block (the mock code repository content).\
                    "cache_control": {"type": "ephemeral"},\
                }\
            ],\
        },\
        {\
            "role": "user",\
            "content": user_input,\
        },\
    ]
    completion = client.chat.completions.create(
        # Select a model that supports explicit cache.
        model="qwen3-coder-plus",
        messages=messages,
    )
    return completion

# First request
first_completion = get_completion("What is the content of this code?")
created_cache_tokens = first_completion.usage.prompt_tokens_details.cache_creation_input_tokens
print(f"First request cache creation tokens: {created_cache_tokens}")
hit_cached_tokens = first_completion.usage.prompt_tokens_details.cached_tokens
print(f"First request cached tokens: {hit_cached_tokens}")
print(f"First request other tokens: {first_completion.usage.prompt_tokens-created_cache_tokens-hit_cached_tokens}")
print("=" * 20)
# Second request. The code content is the same, only the question is different.
second_completion = get_completion("What are some possible optimizations for this code?")
created_cache_tokens = second_completion.usage.prompt_tokens_details.cache_creation_input_tokens
print(f"Second request cache creation tokens: {created_cache_tokens}")
hit_cached_tokens = second_completion.usage.prompt_tokens_details.cached_tokens
print(f"Second request cached tokens: {hit_cached_tokens}")
print(f"Second request other tokens: {second_completion.usage.prompt_tokens-created_cache_tokens-hit_cached_tokens}")

This example caches the code repository content as a prefix for subsequent requests that ask different questions about the same repository.

plaintext
First request cache creation tokens: 1605
First request cached tokens: 0
First request other tokens: 13
====================
Second request cache creation tokens: 0
Second request cached tokens: 1605
Second request other tokens: 15

To ensure model performance, the system appends a few internal tokens. These tokens are billed at the standard input price. For more information, see the FAQ.

Caching the tool list during function calling

When caching system messages in a function calling scenario, the tools parameter is included as part of the system message for caching. You must ensure that the tool definition in each request is identical, including the tool order, field order, and field structure. You must add the cache_control marker to the content block that serves as the end of your cacheable prefix.

The following example shows the complete process: the first request creates a cache, and the second request results in a cache hit.

python
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)

# Mock code repository content, ensuring it exceeds the minimum 1,024-token threshold for explicit cache.
long_text_content = "<Your Code Here>" * 400

# Tool definition: Ensure it is identical for every request (tool order, field order, and field structure).
tools = [\
    {\
        "type": "function",\
        "function": {\
            "name": "get_weather",\
            "description": "Get the current weather information for a specified city.",\
            "parameters": {\
                "type": "object",\
                "properties": {\
                    "city": {\
                        "type": "string",\
                        "description": "The city name, e.g., Beijing, Shanghai, or New York."\
                    },\
                    "unit": {\
                        "type": "string",\
                        "description": "The temperature unit, 'celsius' or 'fahrenheit'. Defaults to 'celsius'.",\
                        "enum": ["celsius", "fahrenheit"]\
                    }\
                },\
                "required": ["city"],\
                "additionalProperties": False\
            },\
            "strict": True\
        }\
    },\
    {\
        "type": "function",\
        "function": {\
            "name": "get_current_time",\
            "description": "Get the current date and time for a specified time zone.",\
            "parameters": {\
                "type": "object",\
                "properties": {\
                    "timezone": {\
                        "type": "string",\
                        "description": "IANA time zone name, e.g., 'Asia/Shanghai' or 'America/New_York'. Defaults to 'Asia/Shanghai'."\
                    }\
                },\
                "required": [],\
                "additionalProperties": False\
            },\
            "strict": True\
        }\
    },\
    {\
        "type": "function",\
        "function": {\
            "name": "convert_currency",\
            "description": "Convert currency amounts based on real-time exchange rates.",\
            "parameters": {\
                "type": "object",\
                "properties": {\
                    "from_currency": {\
                        "type": "string",\
                        "description": "The ISO 4217 code of the source currency, e.g., CNY, USD, or EUR."\
                    },\
                    "to_currency": {\
                        "type": "string",\
                        "description": "The ISO 4217 code of the target currency."\
                    },\
                    "amount": {\
                        "type": "number",\
                        "description": "The amount to be converted."\
                    }\
                },\
                "required": ["from_currency", "to_currency", "amount"],\
                "additionalProperties": False\
            },\
            "strict": True\
        }\
    }\
]

def get_completion(user_input, messages=None):
    if messages is None:
        messages = [\
            {\
                "role": "system",\
                "content": [\
                    {\
                        "type": "text",\
                        "text": long_text_content,\
                        # Place the cache_control marker here. This creates a cache block with all content from the start of the messages array to the current content block.\
                        # The cache_control marker can only be added to the content of messages, not to tools.\
                        "cache_control": {"type": "ephemeral"},\
                    }\
                ],\
            }\
        ]

    messages.append({"role": "user", "content": user_input})

    completion = client.chat.completions.create(
        # Select a model that supports explicit cache.
        model="qwen3.6-plus",
        messages=messages,
        tools=tools,
        # Disable thinking mode.
        extra_body={"enable_thinking": False},
    )
    return completion

# First request: Create cache
print("=== First request (Create cache) ===")
first_completion = get_completion("What's the weather like in Beijing now?")
usage = first_completion.usage
print(f"Prompt Tokens: {usage.prompt_tokens}")
print(f"Cache creation tokens: {usage.prompt_tokens_details.cache_creation_input_tokens}")
print(f"Cached tokens: {usage.prompt_tokens_details.cached_tokens}")
print(f"Model selected tool(s): {[t.function.name for t in first_completion.choices[0].message.tool_calls or []]}")
print()

# Second request: Same system message and a different question, resulting in a cache hit.
print("=== Second request (Hit cache) ===")
messages = [\
    {\
        "role": "system",\
        "content": [\
            {\
                "type": "text",\
                "text": long_text_content,\
                "cache_control": {"type": "ephemeral"},\
            }\
        ],\
    }\
]
second_completion = get_completion("What's the weather like in Shanghai now?", messages=messages)
usage = second_completion.usage
print(f"Prompt Tokens: {usage.prompt_tokens}")
print(f"Cache creation tokens: {usage.prompt_tokens_details.cache_creation_input_tokens}")
print(f"Cached tokens: {usage.prompt_tokens_details.cached_tokens}")
print(f"Model selected tool(s): {[t.function.name for t in second_completion.choices[0].message.tool_calls or []]}")

Running the code produces output similar to the following:

plaintext
=== First request (Create cache) ===
 Prompt Tokens: 2174
 Cache creation tokens: 2156
 Cached tokens: 0
 Model selected tool(s): ['get_weather']

 === Second request (Hit cache) ===
 Prompt Tokens: 2174
 Cache creation tokens: 0
 Cached tokens: 2156
 Model selected tool(s): ['get_weather']

Continuous multi-turn conversation

In a typical multi-turn chat scenario, you can add a cache marker to the last content block in the messages array for each request. Starting from the second turn, each request will both hit and refresh the cache block from the previous turn, and create a new cache block for the current turn.

python
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    # If using a model in the China (Beijing) region, replace the base_url with: https://dashscope.aliyuncs.com/compatible-mode/v1
    base_url="https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1",
)

system_prompt = "You are a witty person." * 400
messages = [{"role": "system", "content": system_prompt}]

def get_completion(messages):
    completion = client.chat.completions.create(
        model="qwen3-coder-plus",
        messages=messages,
    )
    return completion

while True:
    user_input = input("User: ")
    messages.append({"role": "user", "content": [{"type": "text", "text": user_input, "cache_control": {"type": "ephemeral"}}]})
    completion = get_completion(messages)
    print(f"[AI Response] {completion.choices[0].message.content}")
    messages.append(completion.choices[0].message)
    created_cache_tokens = completion.usage.prompt_tokens_details.cache_creation_input_tokens
    hit_cached_tokens = completion.usage.prompt_tokens_details.cached_tokens
    uncached_tokens = completion.usage.prompt_tokens - created_cache_tokens - hit_cached_tokens
    print(f"[Cache Info] Cache creation tokens: {created_cache_tokens}")
    print(f"[Cache Info] Cached tokens: {hit_cached_tokens}")
    print(f"[Cache Info] Other tokens: {uncached_tokens}")

Run the code and interact with the large language model. Each question you ask will hit the cache block created in the previous turn.

Implicit cache

Supported models

China (Beijing)

Singapore

US (Virginia)

Germany (Frankfurt)

Hong Kong (China)

The following models are all in the Chinese mainland deployment scope.

  • Text generation model
    • Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08, qwen3-max, qwen3-max-preview, qwen-max

    • Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen-plus

    • Qwen Flash: qwen-flash

    • Qwen Turbo: qwen-turbo

    • Qwen Coder: qwen3-coder-plus, qwen3-coder-flash

    • DeepSeek: deepseek-v4-pro, deepseek-v4-flash, deepseek-v3.2, deepseek-v3.1, deepseek-v3, deepseek-r1

    • Kimi: kimi-k2.6, kimi-k2.5, kimi-k2-thinking, Moonshot-Kimi-K2-Instruct

    • GLM: glm-5.1, glm-5, glm-4.7, glm-4.6

    • MiniMax: MiniMax-M2.5

  • Visual understanding model
    • Qwen VL: qwen3-vl-plus, qwen3-vl-flash, qwen-vl-max, qwen-vl-plus

The following models are all in the International deployment scope.

  • Text generation model
    • Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08, qwen3-max, qwen3-max-preview, qwen-max

    • Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen-plus

    • Qwen Flash: qwen-flash

    • Qwen Turbo: qwen-turbo

    • Qwen Coder: qwen3-coder-plus, qwen3-coder-flash

    • DeepSeek: deepseek-v4-pro, deepseek-v4-flash, deepseek-v3.2

    • GLM (deployed on Alibaba Cloud Model Studio): glm-5.1

  • Visual understanding model
    • Qwen VL: qwen3-vl-plus, qwen3-vl-flash, qwen-vl-max, qwen-vl-plus

Different deployment scopes support different models.

  • Global deployment scope:
    • Text generation model
      • Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08, qwen3-max

      • Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen-plus

      • Qwen Flash: qwen-flash

      • Qwen Coder: qwen3-coder-plus, qwen3-coder-flash

      • DeepSeek: deepseek-v4-pro, deepseek-v4-flash

      • Kimi (deployed on Alibaba Cloud Model Studio): kimi-k2.5

    • Visual understanding model
      • Qwen VL: qwen3-vl-plus, qwen3-vl-flash
  • US deployment scope:
    • Text generation model
      • Qwen Plus: qwen-plus-us

      • Qwen Flash: qwen-flash-us

    • Visual understanding model
      • Qwen VL: qwen3-vl-flash-us

Different deployment scopes support different models.

  • Global deployment scope:
    • Text generation model
      • Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08, qwen3-max

      • Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26, qwen-plus

      • Qwen Flash: qwen-flash

      • Qwen Coder: qwen3-coder-plus, qwen3-coder-flash

      • DeepSeek: deepseek-v4-pro, deepseek-v4-flash

      • Kimi (deployed on Alibaba Cloud Model Studio): kimi-k2.5

    • Visual understanding model
      • Qwen VL: qwen3-vl-plus, qwen3-vl-flash
  • EU deployment scope:
    • Text generation model

      • Qwen Max: qwen3-max

      • Qwen Plus: qwen-plus

Visual understanding model - Qwen VL: qwen3-vl-plus, qwen3-vl-flash

Different deployment scopes support different models.

  • Global deployment scope:
    • Qwen Max: qwen3.7-max, qwen3.7-max-2026-05-20, qwen3.7-max-2026-06-08

    • Qwen Plus: qwen3.7-plus, qwen3.7-plus-2026-05-26

  • Hong Kong (China) deployment scope:
    • Text generation model
      • Qwen Max: qwen3-max

      • Qwen Plus: qwen-plus

    • Visual understanding model
      • Qwen VL: qwen3-vl-plus

How it works

This feature activates automatically when you send a request to a supported model. The system works as follows:

  1. Find: After receiving a request, the system uses prefix matching to check the cache for a common prefix within the request's messages array.

  2. Decision:

    • If a cache hit occurs, the system uses the cached result to complete the inference.

    • If a cache miss occurs, the system processes the request normally and may store the prompt's prefix in the cache for subsequent requests.

The system periodically clears inactive cached data. The cache hit ratio is not guaranteed. A cache miss can occur even for identical requests because the system ultimately determines the final hit probability.

Note

The minimum number of tokens to trigger implicit caching is approximately 1000 for the qwen3.7-max series and 256 for other models.

Improve the cache hit ratio

The implicit cache works by identifying a common prefix in different requests. To improve the cache hit ratio, place repeating content at the beginning of your prompt and unique content at the end.

  • Text generation model: Assume the system has cached "ABCD". A request for "ABE" might hit the "AB" portion, while a request for "BCD" results in a cache miss.

  • Visual understanding model:

    • To ask multiple questions about the same image or video: Place the image or video before the text to improve the hit ratio.

    • To ask the same question about different images or videos: Place the text before the image or video to improve the hit ratio.

Billing

There is no extra charge for using the implicit cache.

When a request results in a cache hit, the cached input tokens are billed as cached_token at a discounted rate that varies by model. Input tokens that are not served from the cache are billed at the standard input_token rate. Output tokens are billed at the standard rate.

  • For models other than deepseek-v4-pro, the unit price of cached_token is 20% of the standard input_token price.

  • For deepseek-v4-pro, the cached_token price is not 20% of the input_token price. See the Model Studio console for details.

For example, consider a request with 10,000 input tokens where 5,000 tokens result in a cache hit. The cost is calculated as follows:

  • Uncached tokens (5,000) are billed at 100% of the standard rate.

  • Cached tokens (5,000) are billed at 20% of the standard rate.

The total input cost is therefore 60% of what it would be without a cache: (5,000 × 100% + 5,000 × 20%) / 10,000 = 60%.

The number of cached tokens is available in the cached_tokens property of the response.

OpenAI-compatible - Batch (file input) requests are not eligible for cache discounts.

Cache hit examples

Text generation models

Visual understanding models

OpenAI-compatible

DashScope

Anthropic-compatible

When you call a model using an OpenAI-compatible method and an implicit cache hit occurs, the response is similar to the following. The number of cached tokens, reported in usage.prompt_tokens_details.cached_tokens, is included in the total usage.prompt_tokens.

json
{
    "choices": [\
        {\
            "message": {\
                "role": "assistant",\
                "content": "I am a large-scale language model developed by Alibaba Cloud. My name is Qwen."\
            },\
            "finish_reason": "stop",\
            "index": 0,\
            "logprobs": null\
        }\
    ],
    "object": "chat.completion",
    "usage": {
        "prompt_tokens": 3019,
        "completion_tokens": 104,
        "total_tokens": 3123,
        "prompt_tokens_details": {
            "cached_tokens": 2048
        }
    },
    "created": 1735120033,
    "system_fingerprint": null,
    "model": "qwen-plus",
    "id": "chatcmpl-6ada9ed2-7f33-9de2-8bb0-78bd4035025a"
}

When you call a model using the DashScope Python SDK or an HTTP request and an implicit cache hit occurs, the response is similar to the following. The number of cached tokens, reported in usage.prompt_tokens_details.cached_tokens, is included in the total usage.input_tokens.

json
{
    "status_code": 200,
    "request_id": "f3acaa33-e248-97bb-96d5-cbeed34699e1",
    "code": "",
    "message": "",
    "output": {
        "text": null,
        "finish_reason": null,
        "choices": [\
            {\
                "finish_reason": "stop",\
                "message": {\
                    "role": "assistant",\
                    "content": "I am a large-scale language model from Alibaba Cloud. My name is Qwen. I can generate various types of text, such as articles, stories, and poems, and can adapt them based on different scenarios and requirements. Additionally, I can answer various questions and provide help and solutions. If you have any questions or need assistance, feel free to ask, and I will do my best to provide support. Please note that repeating the same content may not yield a more detailed response. It is recommended that you provide more specific information or vary your questions so I can better understand your needs."\
                }\
            }\
        ]
    },
    "usage": {
        "input_tokens": 3019,
        "output_tokens": 101,
        "prompt_tokens_details": {
            "cached_tokens": 2048
        },
        "total_tokens": 3120
    }
}

When you call a model using an Anthropic-compatible method and an implicit cache hit occurs, you can find the number of cached tokens in usage.cache_read_input_tokens. This value is not included in usage.input_tokens but is reported separately.

json
{
    "id": "msg_01XFDUDYJgAACzvnptvVoYEL",
    "type": "message",
    "role": "assistant",
    "content": [\
        {\
            "type": "text",\
            "text": "This content is repeated placeholder text."\
        }\
    ],
    "model": "qwen3-coder-plus",
    "stop_reason": "end_turn",
    "usage": {
        "input_tokens": 82,
        "cache_creation_input_tokens": 0,
        "cache_read_input_tokens": 1536,
        "output_tokens": 14
    }
}

OpenAI-compatible

DashScope

Anthropic-compatible

When you call a model using an OpenAI-compatible method and an implicit cache hit occurs, the response is similar to the following. The number of cached tokens, reported in usage.prompt_tokens_details.cached_tokens, is included in the total usage.prompt_tokens.

json
{
  "id": "chatcmpl-3f3bf7d0-b168-9637-a245-dd0f946c700f",
  "choices": [\
    {\
      "finish_reason": "stop",\
      "index": 0,\
      "logprobs": null,\
      "message": {\
        "content": "This image shows a heartwarming scene of a woman and a dog interacting on a beach. The woman, wearing a plaid shirt, is sitting on the sand and smiling as she interacts with the dog. The dog is a large, light-colored breed wearing a colorful collar, with its front paw raised as if to shake hands or give a high-five to the woman. The background is a vast ocean and sky, with sunlight shining from the right side of the frame, adding a warm and serene atmosphere to the entire scene.",\
        "refusal": null,\
        "role": "assistant",\
        "audio": null,\
        "function_call": null,\
        "tool_calls": null\
      }\
    }\
  ],
  "created": 1744956927,
  "model": "qwen-vl-max",
  "object": "chat.completion",
  "service_tier": null,
  "system_fingerprint": null,
  "usage": {
    "completion_tokens": 93,
    "prompt_tokens": 1316,
    "total_tokens": 1409,
    "completion_tokens_details": null,
    "prompt_tokens_details": {
      "audio_tokens": null,
      "cached_tokens": 1152
    }
  }
}

When you use the DashScope Python SDK or the HTTP method to call a model and trigger the implicit cache, the number of tokens that hit the cache is included in the total input tokens (usage.input_tokens), and the specific viewing location varies by region and model:

  • China (Beijing):
    • For qwen-vl-max and qwen-vl-plus, the count is in usage.prompt_tokens_details.cached_tokens.

    • For qwen3-vl-plus and qwen3-vl-flash, the count is in usage.prompt_tokens_details.cached_tokens.

  • Asia Pacific SE 1 (Singapore): For all models, the count is in usage.cached_tokens.

Models that currently use usage.cached_tokens will be upgraded to use usage.prompt_tokens_details.cached_tokens in the future.

json
{
  "status_code": 200,
  "request_id": "06a8f3bb-d871-9db4-857d-2c6eeac819bc",
  "code": "",
  "message": "",
  "output": {
    "text": null,
    "finish_reason": null,
    "choices": [\
      {\
        "finish_reason": "stop",\
        "message": {\
          "role": "assistant",\
          "content": [\
            {\
              "text": "This image shows a heartwarming scene of a woman and a dog interacting on a beach. The woman, wearing a plaid shirt, is sitting on the sand and smiling as she interacts with the dog. The dog is a large breed wearing a colorful collar, with its front paw raised as if to shake hands or give a high-five to the woman. The background is a vast ocean and sky, with sunlight shining from the right side of the frame, adding a warm and serene atmosphere to the entire scene."\
            }\
          ]\
        }\
      }\
    ]
  },
  "usage": {
    "input_tokens": 1292,
    "output_tokens": 87,
    "input_tokens_details": {
      "text_tokens": 43,
      "image_tokens": 1249
    },
    "total_tokens": 1379,
    "output_tokens_details": {
      "text_tokens": 87
    },
    "image_tokens": 1249,
    "cached_tokens": 1152
  }
}

When you call a visual understanding model using an Anthropic-compatible method and an implicit cache hit occurs, the number of cached tokens is reported in the usage.cache_read_input_tokens field, which is consistent with text generation models.

json
{
  "id": "msg_01XFDUDYJgAACzvnptvVoYEL",
  "type": "message",
  "role": "assistant",
  "content": [\
    {\
      "type": "text",\
      "text": "This image shows a heartwarming scene of a woman and a dog interacting on a beach."\
    }\
  ],
  "model": "qwen-vl-max",
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 369,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 896,
    "output_tokens": 28
  }
}

Use cases

If your requests share a common prefix, context cache can significantly improve inference speed, lower inference cost, and reduce first-packet latency. This feature is particularly useful in the following scenarios:

  1. Long-text question and answer

Use this pattern when you send multiple requests about the same long text, such as a novel, textbook, or legal document.

Messages array for the first request

python
messages = [{"role": "system","content": "You are a language teacher who can help students with reading comprehension."},\
             {"role": "user","content": "\
```\
\
\
\
**Message array in the subsequent request**\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
```python\
messages = [{"role": "system","content": "You are a language arts teacher. You can help students with reading comprehension."},\
             {"role": "user","content": "<Article content> Please analyze the third paragraph of this text."}]\
```\
\
\
\
Although the questions differ, they are all based on the same article. The identical system prompt and article content create a large amount of repetitive prefix information, making a cache hit highly likely.\
\
2. **Automatic code completion**\
\
In automatic code completion scenarios, the model uses the code in the current context to generate subsequent code. As you continue to write code, the prefix often remains the same, allowing `context cache` to reuse it and accelerate completions.\
\
3. **Multi-turn conversation**\
\
For a `multi-turn conversation`, you append each turn to the `messages` array. This ensures each new request shares a common prefix with the previous turns, increasing the likelihood of a cache hit.\
\
**Messages array for the first turn**\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
```python\
messages=[{"role": "system","content": "You are a helpful assistant."},\
             {"role": "user","content": "Who are you?"}]\
```\
\
\
\
**Messages array for the second turn**\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
```python\
messages=[{"role": "system","content": "You are a helpful assistant."},\
             {"role": "user","content": "Who are you?"},\
             {"role": "assistant","content": "I am Qwen, developed by Alibaba Cloud."},\
             {"role": "user","content": "What can you do?"}]\
```\
\
\
\
As the conversation grows, the benefits of caching for inference speed and cost become more significant.\
\
4. **Role-playing or few-shot learning**\
\
In role-playing or few-shot learning scenarios, you often include extensive guidance in the `prompt` to control the output format. This creates a large shared prefix across requests.\
\
For example, when instructing the model to act as a marketing expert, the `system prompt` contains extensive text. The following are two example requests:\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
```python\
system_prompt = """You are an experienced marketing expert. Please provide detailed marketing suggestions for different products in the following format:\
\
1. Target audience: xxx\
\
2. Main selling points: xxx\
\
3. Marketing channels: xxx\
...\
12. Long-term development strategy: xxx\
\
Please ensure your suggestions are specific, actionable, and highly relevant to the product features."""\
\
# User message for the first request, asking about a smartwatch\
messages_1=[\
{"role": "system", "content": system_prompt},\
{"role": "user", "content": "Please provide marketing suggestions for a newly launched smartwatch."}\
]\
\
# User message for the second request, asking about a laptop. Since the system_prompt is the same, there is a high probability of hitting the cache.\
messages_2=[\
{"role": "system", "content": system_prompt},\
{"role": "user", "content": "Please provide marketing suggestions for a newly launched laptop."}\
]\
```\
\
With `context cache`, the system can respond faster because the lengthy `system prompt` is cached, even when users frequently change the product in their request (for example, from a smartwatch to a laptop).\
\
5. **Video understanding**\
\
In video understanding scenarios, if you ask multiple questions about the same video, placing the `video` before the `text` increases the likelihood of a cache hit. If you ask the same question about different videos, placing the `text` before the `video` increases the likelihood of a cache hit. The following are two example requests for the same video:\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
\
````python\
# User message for the first request, asking about the content of this video\
messages1 = [\
       {"role":"system","content":[{"text": "You are a helpful assistant."}]},\
       {"role": "user",\
           "content": [\
               {"video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250328/eepdcq/phase_change_480p.mov"},\
               {"text": "What is the content of this video?"}\
           ]\
       }\
]\
\
# For the second request about the same video, placing the video before the text increases the likelihood of a cache hit.\
messages2 = [\
       {"role":"system","content":[{"text": "You are a helpful assistant."}]},\
       {"role": "user",\
           "content": [\
               {"video": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20250328/eepdcq/phase_change_480p.mov"},\
               {"text": "Please describe the series of events in the video. Output the start time (start_time), end time (end_time), and event (event) in JSON format. Do not output the ```json``` code block."}\
           ]\
       }\
]\
````\
\
\
## **FAQ**\
\
### **Q: How do I disable implicit cache?**\
\
A: You cannot disable it. The implicit cache is enabled for all applicable model requests as long as it does not affect response quality. When a cache hit occurs, it reduces costs and improves response speed.\
\
### **Q: Why did my explicit cache miss?**\
\
A: A cache miss occurs for the following reasons:\
\
- The system clears the cache block if it is not hit within its 5-minute validity period.\
\
- A cache miss occurs if there are more than 20 `content` blocks between the last `content` in the prompt and an existing cache block. We recommend creating a new cache block.\
\
\
### **Q: Does a cache hit reset its validity period?**\
\
A: Yes. Each hit resets the cache block's validity period to 5 minutes.\
\
### **Q: Is explicit cache shared between accounts?**\
\
A: No. Both implicit and explicit cache data are isolated at the account level.\
\
### **Q:** Is explicit cache shared across models?\
\
A: No. Cache data is isolated between models.\
\
### **Q: Why isn't**`input_tokens` **in**`usage` **the sum of**`cache_creation_input_tokens` **and**`cached_tokens` **?**\
\
A: To ensure the quality of the model output, the backend service appends a small number of tokens (typically fewer than 10) to the prompt that you provide. These tokens are placed after the `cache_control` marker. Therefore, they are not counted toward cache creation or reading, but they are included in the total `input_tokens`.\
\
Previous: Partial modeNext: Batch inference\
\
Is this page helpful?\
\

\

\
Explicit cache\
\
How it works\
\
Supported models\
\
Quick start\
\
Fine-grained cache control\
\
Billing\
\
Cacheable content\
\
Limitations\
\
Function calling cache optimization\
\
Usage examples\
\
Implicit cache\
\
Supported models\
\
How it works\
\
Improve the cache hit ratio\
\
Billing\
\
Cache hit examples\
\
Use cases\
\
FAQ\
\
Q: How do I disable implicit cache?\
\
Q: Why did my explicit cache miss?\
\
Q: Does a cache hit reset its validity period?\
\
Q: Is explicit cache shared between accounts?\
\
Q: Is explicit cache shared across models?\
\
Q: Why isn't input\_tokens in usage the sum of cache\_creation\_input\_tokens and cached\_tokens?\
\
\
\
 [](https://www.alibabacloud.com/help/en/model-studio/context-cache#top "Back to Top")\
\
Contact Us\
\
\
\
## Sales Support\
\
Live-chat with our sales team or get in touch with a business development professional in your region.\
\
Contact Sales\
\
## Technical Support\
\
Open a ticket and get quick help from our technical team.\
\
Open a Ticket >\
\
## Connect & Report Abuse\
\
We look forward to your suggestion.\
\
Post a Suggestion > Report Abuse >\
\
Chat now with **Alibaba Cloud Customer Service** to assist you in finding the right products and services to meet your needs.\
\
 [](https://www.alibabacloud.com/ccim/mobile) [](https://int.alibabacloud.com/m/1000409525/) [](https://wa.me/8613811961093)\
\
\\
\\
Hi, I'm Alibaba Cloud AI Assistant!\\
\\
I can help with questions and solutions.\
\
Why Alibaba Cloud\
\
About Alibaba Cloud\
\
Asia Accelerator\
\
Our Global Network\
\
Global Offices\
\
Trust Center\
\
Case Studies\
\
Analyst Reports\
\
Products & Pricings\
\
Pricing Calculator\
\
ECS\
\
SAS\
\
Model Studio\
\
Database\
\
Security\
\
SMS\
\
Solutions\
\
Financial Services\
\
Retail Services\
\
Media Services\
\
Gaming Services\
\
ISV Solutions\
\
Engage\
\
Developer Community\
\
Partner Network\
\
Startups\
\
Marketplace\
\
Join Alibaba Cloud\
\
Resources & Support\
\
Developer Learning Hub\
\
Documentation Center\
\
Training & Certification\
\
Service Notices\
\
Submit a Ticket\
\
Security Report\
\
Qwen Cloud\
\
Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links\
\
© 2009-2026 Copyright by Alibaba Cloud All rights reserved\
\
- Facebook\
- Linkedin\
- Twitter\
- YouTube\
- TikTok\
- contact.us@alibabacloud.com\
- Call Us Now\
- Discord\
\
[](https://privacy.truste.com/privacy-seal/validation?rid=e83a96cf-eaa2-4c82-8d85-883196fcb42c)\
\
[](https://privacy.truste.com/privacy-seal/validation?rid=51be162a-c152-402b-84d2-d0a724de679a)\
\
© 2009-2026 Copyright by Alibaba Cloud All rights reserved\
\
Careers About Us Privacy Policy Legal Integrity Compliance Reporting Channel Service Notices Links\
\
```\
\