Alibaba Cloud’s AI story has moved beyond Qwen as a standalone model. The larger question is how cloud infrastructure itself must be rewritten when the main user shifts from human developers to agents.
This essay is part of China Industry Signals, a high-frequency rolling series tracking the latest trends of systemically important companies, sectors, technologies, and supply chains across China’s industrial system.
Three questions this essay tracks
First, what does it mean to make cloud products agent-ready?
Alibaba Cloud is turning model services into Skills and CLI tools, making Qwen Cloud less like a traditional model marketplace and more like a tool layer for agents.
Second, why does agent workload change cloud pricing and scheduling?
QoderWork’s peak-valley Token shows that long-running agents can be scheduled like workloads, not used like chatbots. Night-time execution, discounted tokens, and asynchronous task completion may become part of AI cloud pricing.
Third, can Alibaba turn cloud operations themselves into agent workloads?
AliyunConsoleAgent, Agent Security Center, and Agent Firewall point to a deeper direction: the cloud is not only hosting agents; the cloud itself is becoming managed, checked, and protected by agents.
Clear Dawn over Streams and Mountains (溪山清晓), by Xie Zhiliu (谢稚柳)
Its quiet morning atmosphere, layered mountain paths, flowing streams, and carefully ordered spatial structure echo the central argument of this essay: Alibaba Cloud is not only building stronger models, but reorganizing cloud infrastructure into a system that agents can understand, call, schedule, and govern.
In the past, the way developers used cloud platforms was straightforward.
A developer opened the console, read documentation, selected a model, compared prices, copied an API key, configured the environment, checked the bill, set permissions, and handled security risks. If he wanted to deploy an application, he had to move back and forth among models, databases, object storage, networking, security, monitoring, logs, function compute, containers, and servers.
This system was designed for humans.
Pages, buttons, documents, consoles, SDKs, APIs, permission menus, and billing systems assumed that the default users were developers, operations teams, and enterprise IT departments. Cloud vendors provided resources and products. Human users understood, selected, configured, and operated them.
The agent era changes that assumption.
An agent does not need to browse pages like a human. It needs structured capabilities, clear interfaces, callable tools, task templates, command lines, state feedback, permission boundaries, and audit records. It will not only chat with users in real time during the day; it may also queue tasks at night. It will not only use cloud resources; it may also check cloud documents, validate consoles, configure cloud products, discover risks, and trigger security policies.
This is the common direction behind Alibaba Cloud’s recent moves.
On May 20, Alibaba Cloud announced Qwen Cloud, a full-stack “chip–cloud–model–inference” agentic upgrade, the self-developed Zhenwu M890 AI chip and supernode servers, and Qwen3.7-Max. In June, QoderWork launched peak-valley Token pricing. After that, the AliyunConsoleAgent paper and Alibaba Cloud documentation on Agent Security Center and Agent Firewall pushed the same direction further.
This shows that Alibaba Cloud is working on something disruptive: turning the cloud from human-operated infrastructure into agent-operated infrastructure.
This may become one of the most important questions in the future of cloud computing: if large numbers of tasks are executed by agents, how should cloud platforms turn models, cloud products, inference resources, pricing systems, consoles, and security systems into systems that agents can understand, call, schedule, and govern?
Qwen Cloud: The model platform begins to become an agent tool layer
Qwen Cloud is the first entry point for understanding Alibaba’s Agentic Cloud.
On May 20, Alibaba Cloud launched “Qwen Cloud,” positioning it as an AI product website built for the agent era. It provides APIs for more than 150 mainstream models, including Qwen, GLM, Kimi, DeepSeek, Wan, and HappyHorse, and packages the core capabilities of model services into Skills and CLI tools, allowing agent tools to use models and develop AI applications more efficiently.
In the past, the core users of model-service platforms were humans. Developers opened a webpage, looked at a model list, compared context length, price, modality, parameters, capability descriptions, and applicable scenarios, then copied the API and wrote code to integrate it. That workflow works for human developers, but it is not friendly to agents.
Agents need callable model capabilities. They need scriptable and automatable interfaces. They need clear information about model differences. They need workflows that can be executed through CLI and Skills. By turning model services into Skills and CLI tools, Qwen Cloud expands the default user of the model-service platform from humans to agents.
This will change the shape of model platforms.
A model marketplace built for humans focuses on displaying models. A model platform built for agents focuses on turning models into tools. Human developers need page navigation, documentation, and trial interfaces; agents need structured calls, parameter descriptions, capability boundaries, price information, task adaptation, and automated execution paths.
Qwen Cloud provides comparison information across parameters, capability, price, context length, modality support, and applicable scenarios during model selection and model calling. Users can enter model experience pages directly after comparison and test outputs with real prompts or tasks.
These functions are useful for humans, and also useful for agents.
If an agent has to complete a complex task, it may need to select different models: one model for code, one for image, one for low-cost batch processing, one for long context, one for complex reasoning, and one for video generation. If a model platform can structure model capability, price, context, modality, and use cases, it provides the basic information for agent routing.
This is an extension of the Qwen Coding Plan discussed in the previous AI Coding essay.
Qwen Coding Plan puts Qwen, GLM, Kimi, MiniMax and other models into a fixed monthly plan and makes them compatible with multiple AI programming tools. Qwen Cloud goes further by packaging the core capabilities of model services into Skills and CLI. The former solves how developers use multiple models at more controllable cost. The latter solves how agents use models as callable tools.
Alibaba Cloud’s Skills Portal documentation points in the same direction.
The Skills Portal provides standardized skill collections for AI agents, supports natural-language operation of cloud resources, cross-product task orchestration, and security constraints, and is compatible with mainstream agent clients such as Cursor, Claude Code, Qwen Code, Qoder, Codex, Gemini CLI, GitHub Copilot, and OpenClaw. Agents can automatically call relevant Skills through natural-language instructions. The Skills Portal itself is free, but cloud resources created or used by agents through Skills are still billed according to the corresponding cloud products.
This design matters.
It turns Alibaba Cloud products from “buttons and documents inside a human console” into capability packages that agents can install and call. In the future, a developer may not need to read documentation product by product, or manually configure resources in the console. An agent can call elastic compute, databases, object storage, cloud networking, security, function compute and other capabilities through Skills, completing more complex cross-product task orchestration.
This also gives Alibaba Cloud a new platform position.
If more agent clients call cloud products through Alibaba Cloud Skills, Alibaba Cloud is not only providing underlying resources; it is also providing the tool layer for the agent era. Models, Skills, CLI, MCP, billing, permissions, and cloud products will gradually connect into a new developer entry point.
Qwen Cloud can therefore be understood as the first layer of Alibaba’s Agentic Cloud: model inventory, model comparison, Skills, CLI, multi-model calling, and agent tool entry.
“Chip–cloud–model–inference”: Alibaba Cloud’s full-stack Agentic Cloud
Qwen Cloud solves model services and agent calling. Alibaba Cloud’s larger narrative is the full-stack agentic upgrade across “chip–cloud–model–inference.”
At the May 20 launch, Alibaba Cloud announced the completion of its “chip–cloud–model–inference” full-stack agentic upgrade, while releasing Qwen Cloud, the self-developed Zhenwu M890 AI chip, supernode servers equipped with the chip, and Qwen3.7-Max. Alibaba Cloud senior vice president Liu Weiguang said that after agents cross the critical point, they can work 24 hours a day, creating enormous demand for AI and cloud. Alibaba Cloud is upgrading from underlying chips, Agentic Cloud, models, to inference platforms, building “China’s largest AI factory.”
This framing is very Alibaba.
Alibaba is not a pure model company. It is also not a pure AI application company. It has the Qwen model family, Alibaba Cloud, Model Studio, Qwen Cloud, Qoder / QoderWork, T-Head chips, databases, cloud security, compute resources, DingTalk, Taobao, Amap, Ele.me, Fliggy, and a large base of enterprise customers. Its most natural story is a complete AI stack.
Qwen3.7-Max is the model layer.
Qwen Cloud and Model Studio are the model-service layer.
Qoder, QoderWork, Qwen Code, Skills, and CLI are the development and agent-tool layer.
Alibaba Cloud is the compute, storage, network, database, security, and enterprise-customer layer.
Zhenwu M890, supernode servers, and T-Head are the underlying hardware and dedicated AI infrastructure layer.
This combination makes Alibaba Cloud’s AI path different from ByteDance’s and Tencent’s.
ByteDance looks more like AI usage / revenue conversion: Doubao brings consumer usage, Volcano Ark captures MaaS calls, and Seedance enters content-production budgets. Tencent looks more like a service routing layer: WeChat, mini-programs, payments, social relationships, A2A, and service transactions. Alibaba Cloud looks more like infrastructure conversion: turning models, developer tools, cloud resources, inference platforms, hardware, and security systems into infrastructure that agents can run on, call, and govern.
Agent workloads increase infrastructure requirements. A normal chatbot call is relatively simple: the user inputs, the model outputs. A long-running agent may call multiple models, tools, databases, file systems, browsers, code environments, logging systems, cloud resources, and security policies. It may run continuously for hours or longer. It may fail, retry, split tasks, call sub-agents, write files, access external services, trigger deployments, and incur resource costs.
This workload places more complex demands on the cloud.
It needs stable inference, tool permissions, context storage, task-state management, cost control, security isolation, log auditing, asynchronous scheduling, and high utilization of underlying compute. If traditional cloud products remain designed only around human clicks and API calls, they will struggle to support large-scale agent workflows.
That makes Alibaba Cloud’s full-stack “Agentic Cloud” narrative a natural choice.
It wants to repackage cloud resources, model services, and development tools into a system that agents can use. The cloud will no longer only be a place that hosts applications. It will also become the environment where agents execute tasks, call tools, manage resources, generate artifacts, and accept governance.
This also has capital-market significance.
Alibaba needs to prove that its core position in China’s AI era comes not only from Qwen model capability, but also from cloud infrastructure. Cloud is one of Alibaba’s strategic assets. If the agent era arrives, cloud platforms will sell not only compute, but also model calls, agent runtime, Skills, task scheduling, security governance, enterprise tools, and industry solutions.
Alibaba Cloud therefore wants to define itself as China’s most complete agent infrastructure company.
Whether that positioning can hold depends on several questions: whether Qwen models can remain competitive; whether Qwen Cloud and Model Studio can become model-service entry points for enterprises and agents; whether Qoder / QoderWork can enter developer workflows; whether Alibaba Cloud can turn large numbers of products into Skills; whether Zhenwu M890 and supernodes can improve AI workload costs; and whether enterprise customers are willing to deploy agent workflows on Alibaba Cloud.
These questions are still being tested.
But Alibaba Cloud’s direction is clear: it does not want to be only a model provider, and it does not want to be only a cloud-resource provider. It wants to become the infrastructure platform of the agent era.
Peak-valley Token: Agent workloads begin to change cloud pricing
QoderWork’s peak-valley Token is one of the most interesting recent signals.
On June 24, Alibaba Cloud announced that QoderWork had launched “peak-valley Token.” Users running agents between 22:00 and 08:00 can automatically receive night-time discounts. Qwen3.7-Max can be discounted to as low as 20% of the daytime price. The night-time discount covers QoderWork, Qoder Desktop, CLI and other products. Users can set scheduled tasks during the day, or submit long-horizon task instructions before sleeping, letting agents autonomously execute the full workflow at night and review the results the next morning. Credit consumption is only 20%–40% of the daytime level.
Normal chat models emphasize real-time interaction. The user asks, and the model answers. Latency, response speed, and concurrency experience matter greatly. Users will not accept waiting until the next morning for an ordinary chat answer.
Agent workflows are different.
Many agent tasks do not need to be completed in real time. Code refactoring, test repair, batch data processing, web-page organization, report generation, long-document reading, asset production, spreadsheet cleaning, first drafts of presentations, competitor analysis, log analysis, and task inspection can all be queued. Users can assign tasks during the day, run them at night, and inspect results the next morning. For many enterprises and developers, as long as the result is usable, real-time response is not the first requirement.
This will change AI inference pricing.
In the past, AI model pricing was centered on tokens per million, context length, input and output prices, cache prices, and model tiers. Agent workloads add new dimensions: execution time window, task priority, asynchronous execution, failure retry, tool calling, resource occupation, result validation, and SLA.
Peak-valley Token shows that Alibaba Cloud has begun to treat agents as schedulable workloads.
Traditional cloud computing has long used similar logic. Night-time batch processing, low-priority tasks, reserved instances, spot instances, elastic scheduling, hot and cold storage are all ways to use differences in resource utilization. AI inference will face similar issues. Daytime interactive peaks need low-latency resources. At night, if compute resources are idle, discounts can attract long tasks and low-priority tasks.
This has direct value for cloud vendors.
It can improve resource utilization, smooth load peaks and troughs, reduce user price sensitivity, and turn agent tasks from one-time calls into long-running workflows. For users, peak-valley Token lowers long-task costs. For Alibaba Cloud, it brings more agent workloads onto its platform and locks user habits into QoderWork, Qoder Desktop, and CLI.
This design also shows that AI cloud will increasingly resemble a task-scheduling system.
The unit of a chatbot is a conversation. The unit of an agent is a task. A task can have a start time, deadline, priority, budget, model choice, tool permissions, failure retry policy, and acceptance criteria. If cloud platforms can manage these dimensions, they can move from simply selling tokens to selling task-execution capability.
Peak-valley Token is only the first step.
More pricing models may appear: low-priority agent tasks, night-time batch agents, SLA-tiered pricing, result-based pricing, tool-call-based pricing, spot inference, agent workflow packages, and enterprise task queues. Real-time chat will still exist, but many enterprise and developer workloads will look more like asynchronous computing.
This is especially important for Alibaba Cloud.
Alibaba Cloud is already a cloud vendor, with resource scheduling, billing, task queues, storage, databases, logs, security, and enterprise account systems. The more complex agent workflows become, the more they need cloud-platform scheduling and governance capability. Model companies can provide models. Tool companies can provide interfaces. But cloud platforms can connect compute, tasks, billing, and enterprise accounts.
QoderWork’s peak-valley Token makes this direction clear: agents are not only intelligent assistants inside chat interfaces; they are also a new type of workload inside cloud resource scheduling and billing systems.
AliyunConsoleAgent: The cloud itself will be managed by agents
A deeper layer of Alibaba’s Agentic Cloud is letting agents manage the cloud itself.
The AliyunConsoleAgent paper provides a useful example.
Large cloud platforms have hundreds of products. Console UI changes rapidly, and documentation workflows often become inconsistent with the actual console. A cloud product document that works today may fail months later because button positions, permission logic, product parameters, or console interactions have changed. To verify whether the operation flow in documentation can still be executed end to end in the current console, continuous inspection is required.
The paper estimates that this type of document and console-flow validation requires about 4 million recurring inspections per year, while manual coverage is below 1%.
This number captures the complexity of cloud platforms themselves.
Cloud vendors do not only provide compute and models to customers. They also have large amounts of documentation, consoles, product configurations, permission rules, and operation flows to maintain. Traditional approaches rely on manual testing, documentation engineers, product managers, and customer feedback. As product counts rise and iteration accelerates, manual inspection cannot cover enough.
AliyunConsoleAgent aims to automatically validate documentation workflows inside real cloud-console environments. It has to execute concrete tasks, handle real product states, real permissions, real resource creation, real backend audit logs, and real environmental noise.
The paper proposes a two-stage training method: first using frontier-model trajectory distillation for supervised fine-tuning, then using GRPO reinforcement learning inside real cloud environments, with a dual-channel outcome reward model. To support large-scale RL, the system uses Terraform-based resource pre-configuration and LLM-driven on-demand configuration, reducing environmental noise in the training signal.
This technical detail matters.
The difficulty of cloud-console tasks is not only understanding a page. The agent has to understand cloud product logic, documentation steps, permission conditions, button changes, resource states, and backend results. Whether a task succeeds should not be judged only by the model saying it is complete. It should be objectively evaluated through backend audit logs. The paper’s rule-based reward evaluation protocol is designed to solve reward hacking and result-verification problems.
The experimental results are also meaningful.
On a benchmark containing 278 tasks, the best frontier model achieved a success rate of 65.34%. AliyunConsoleAgent-32B achieved an average success rate of 63.52%, improving over the base model by 20.24 percentage points, narrowing the gap with the best closed frontier model to 1.82 percentage points, while reducing inference cost by 92%.
These figures point to one direction: a specialized cloud-console agent can approach frontier closed-model performance in real cloud environments while sharply reducing cost.
This matters for cloud vendors.
If 4 million annual documentation and console-flow inspections all depend on the most expensive closed models, cost and privacy both become problems. If a specialized agent can be trained and run inside the company’s own cloud environment, cost is much lower and data remains internal. This pattern can extend to more cloud operations tasks: resource configuration, documentation validation, permission inspection, troubleshooting, cost optimization, security response, console automated testing, and customer support.
This shows a deeper meaning of Agentic Cloud.
Cloud platforms do not only host agents. Their own maintenance, verification, operations, and governance will also become agentic. If Alibaba Cloud can turn this capability into an internal productivity tool and then gradually open it to enterprise customers, it will create a new cloud product category.
In the future, enterprises may not only buy servers and model APIs. They may also buy cloud operations agents, security inspection agents, cost optimization agents, documentation validation agents, compliance audit agents, and console automation agents.
This will change the value structure of cloud products.
In the past, cloud vendors sold resources and managed services. Now, cloud vendors can sell “agents that know how to operate the cloud.” Customers need not only compute, but also intelligent agents that can help configure, inspect, optimize, protect, and explain cloud resources. The more these agents understand cloud products, the more valuable they become.
AliyunConsoleAgent is an early signal of Alibaba Cloud turning the cloud itself into an agentic system.
Agent Security Center and Agent Firewall: Agents also need governance
Agents can call tools, access models, connect to external services, read knowledge bases, execute commands, create resources, and access networks. The more capable they become, the larger the risk becomes.
Alibaba Cloud’s recent documentation on Agent Security Center and Agent Firewall makes this problem concrete.
Agent Security Center focuses on security risks after agents are deployed at scale, including prompt injection, jailbreak attacks, obfuscation smuggling, privacy leakage, and Skills poisoning. Skills poisoning deserves particular attention: attackers can implant malicious or hidden instructions inside Skill files loaded by agents, inducing agents to execute unintended actions under the user’s identity when calling that Skill, causing sensitive information leakage, data destruction, or permission abuse.
These risks differ from traditional cloud security.
Traditional cloud security focuses on server vulnerabilities, network boundaries, identity permissions, data leakage, malicious processes, and configuration errors. Agent security also has to cover prompts, tools, models, context, Skills, MCP Servers, external APIs, and autonomous execution behavior. An agent may not have a traditional malicious process, but it can still execute dangerous actions because it was prompt-injected or loaded a malicious Skill.
Agent Firewall addresses these problems from the perspective of network traffic and behavior control.
Alibaba Cloud documentation says Agent Firewall can automatically identify agent instances and the MCP tools and LLM models they call, show egress IPs, associated workloads, and agent counts, and manage runtime network-traffic risks for agents. It supports limiting agent external connections based on domain blacklists, ports, and applications; it can parse traffic and restore original Skill files for cloud sandbox threat detection; it can also monitor communication content and identify AK/SK, API keys, personal information, and other sensitive data exfiltration.
These capabilities directly match the core concerns of enterprise agent deployment.
Enterprises are not only worried that agents will answer incorrectly. They are more worried that agents may access unauthorized model services, connect to the external internet, leak API keys, send customer data to external LLMs, load poisoned Skills, execute dangerous commands, bypass permission boundaries, or exfiltrate sensitive files.
The significance of Agent Firewall and Agent Security Center is that they turn agents into cloud-security assets.
In the past, cloud security centers identified servers, containers, databases, load balancers, elastic public IPs, accounts, and permissions. Now they also need to identify agent instances, models called by agents, tools called by agents, external connection paths, Skills loaded by agents, communication logs between agents and LLMs, and agent data-exfiltration risks.
This will become a precondition for enterprise agent deployment.
Without security governance, enterprises will find it difficult to allow agents to autonomously call tools, access knowledge bases, and execute tasks. The more useful an agent is, the more permissions it needs; the greater the permissions, the more audit and control are required. Enterprises will not only ask whether the model is smart. They will also ask whether agents can be discovered, limited, audited, blocked, and rolled back.
Alibaba Cloud’s advantage here is that it already has cloud-security product lines, enterprise customers, network boundaries, and permission systems. Agent security can naturally connect with Cloud Firewall, Security Center, SASE, AI security guardrails, and OpenClaw real-time protection.
This also shows that the next layer of Agentic Cloud will include governance.
Agents do not only need to execute tasks. They also need to be managed. Enterprises need to know which agents are running, which models they can call, which MCP tools they connect to, which external services they access, whether they leak sensitive information, whether they load abnormal Skills, and whether they execute dangerous commands.
Cloud vendors that can provide this governance system will gain a key position in enterprise agent deployment.
Alibaba Cloud is therefore not only selling models to developers. It is turning cloud infrastructure into a system that agents can understand, call, schedule, and govern.
Implications
First, cloud-platform competition will move from “selling resources” to selling agent-callable capabilities.


