Cloud Infrastructure · CIPU / DPU

What the DPU Is
in the Agentic Era

How CIPU builds Alibaba Cloud’s core technical competitiveness in AI infrastructure — for AI agents, inference, training, and beyond. From multi-tenant secure isolation and elastic bare metal to agent sandboxes, PD-disaggregated inference and lossy multi-path training networks.

CI TECHNICAL ESSAY·12 MIN READ·CIPU / DPU / AGENT INFRA

TL;DR

Recently, Neo Clouds have exposed many security problems, and cases have emerged such as sandbox vulnerabilities at OpenAI and others being exploited by models to attack Hugging Face. As LLMs grow ever more capable, Nvidia has also begun promoting Scale-In, which aims to transform the traditional north-south network into a unified infrastructure domain purpose-built to provide security, management, and data acceleration for Agentic AI Factories.

So what is the DPU in the Agentic era? In fact, the answer has been around for a long time. The term "DPU" is poorly defined; it originally came from Fungible, who never quite figured it out, and Nvidia, which inherited the term, hasn't figured it out either. If CPUs and GPUs can process data, should they be called DPUs too? If switches and even FPGAs can also process data, do they count as DPUs as well? In fact, we defined it clearly long ago: the correct term should be the Cloud Infrastructure Processing Unit (CIPU). The "Scale-In DPU" framing is essentially an unfolding of this very concept, further clarifying what a DPU is.

At this year's Apsara Conference, we presented a dedicated session, "An In-Depth Technical Analysis of the Cloud Infrastructure Processing Unit (CIPU)"[1], which elaborates in detail on how CIPU builds Alibaba Cloud's core technical competitiveness in AI infrastructure for AI Agent, inference, training, and more.

1. Reaffirming the Core Value Proposition of Cloud Computing in the AI Era

Because compute resources remain persistently scarce, we often see all kinds of small ads on various social media: "xx 8-GPU servers available for rent at xx thousand dollar, xx units currently in stock." We can generally call this kind the social-media cloud; on the other hand, there is the Neo Cloud represented by the likes of CoreWeave. The entire industry seems to be frantically expanding compute, with all the attention on performance. So what distinguishes these clouds from traditional cloud service providers (AWS, GCP, Azure, Alibaba Cloud, etc.)?

For this reason, we need to discuss here in detail: What is the core value of cloud computing in the AI era?

Figure 1
Fig. 1

Let's first talk about the core business value of computing. You certainly don't want your own code to be leaked by mistake, or your core business to incur financial losses due to stability issues. But as models grow ever more capable and the Agentic era arrives, a series of problems have emerged... You certainly don't want your code to be wrongly uploaded and leaked, nor do you want some core system to be mistakenly attacked and fail. Therefore, security and stability are the cornerstones of computing. Only then come performance and cost: can the same processor deliver 10%–20% higher performance than other compute providers, or the same performance at 10%–20% lower cost? These value propositions must be considered whether for cloud computing or for on-premises (localized) deployment.

So what is cloud computing? What is the core business definition of cloud computing? It focuses mainly on elasticity and multi-tenancy. How should we understand these two terms? In fact, because of the large amount of IDC construction and private (on-premises) deployment in China, many people have an insufficient understanding of elasticity—especially in the Agentic era, where certain tasks may temporarily require millions of CPU cores of compute. The so-called elasticity is essentially injecting compute liquidity into these compute demands. The cloud service provider thereby effectively becomes a compute financial institution, particularly in recent years when compute has been scarce. Since the cloud is providing compute liquidity for multiple customers, secure isolation between tenants becomes the single most important bottom line—just as you don't want the money you keep in a bank to be stolen by others, for computing you likewise don't want your critical data to be accessed/tampered with or even deleted by others... I previously wrote a detailed article on this topic, "A Sharp Critique of a Competitor's Claim That Traditional Clouds Are Still Just Selling Hardware: Discussing Cloud Computing and Its Liquidity Management from a Financial Perspective", which is worth reading in full.

Therefore, in the AI era, we reaffirm the core value proposition of cloud computing: Multi-tenant secure isolation remains the bottom line: without multi-tenant secure isolation, all other business value drops to zero. It is precisely these two core business definitions of cloud computing—elasticity and multi-tenancy—that lead to the conclusion that cloud infrastructure needs a new processor to be responsible for handling these tasks. This is also what the concept of the DPU failed to articulate over the past few years, and the reason Scale-In is now being reaffirmed. We defined this class of processor long ago: the Cloud Infrastructure Processing Unit: CIPU. Its primary business positioning is to build enhanced multi-tenant secure isolation (security) + hardware acceleration of data access (performance).

Figure 2
Fig. 2

CIPU's development has undergone exactly 10 years of accumulation and five generations of architectural evolution, but the one thing that has never changed is that from day one we have always centered on the core business value of cloud computing.

Figure 3
Fig. 3

First is the elastic bare-metal instance, a technology pioneered by the Alibaba Cloud CIPU team and released in October 2017. Even AWS, the number-one player in cloud computing, was a month behind us. And today, whether it's GPU compute servers or the CPU servers used for the Agent Sandbox, all are built on top of bare-metal technology... Then the second-generation CIPU achieved the pooling of elastic bare-metal and VMs together—in financial terms, securitization of compute to provide better compute liquidity. Next, the third generation implemented hardware acceleration on the VPC network, and the fourth generation further built storage / local-disk virtualization acceleration capabilities, and laid out, very early on, elastic RDMA (eRDMA) capabilities targeting the cloud's needs for elasticity and multi-tenant isolation—becoming the world's first cloud service provider to offer the standard RDMA RC Verbs interface on all CPU/GPU instances across all Regions and all AZs. The fifth-generation CIPU further added acceleration for cloud parallel file systems such as CPFS, as well as data encryption and TPM trusted-computing capabilities.

2. What New Business Opportunities and Technical Challenges Does Cloud Infrastructure Face in the AI Era?

Over the past few months, as LLMs have grown ever more capable, a large number of security risks have been exposed. First, in July 2026, because an agent escaped from its sandbox, OpenAI launched an attack on Hugging Face; then the media gradually began to pay attention to infrastructure security. On August 30, 2026, SemiAnalysis published a dedicated article revealing that most Neo Clouds are riddled with security vulnerabilities, "Most Neoclouds Suck At Security"[2]. It also drew the attention of many industry heavyweights.

Figure 4
Fig. 4

For a more detailed introduction, see "On the Security Issues of Neo Clouds".

Figure 5
Fig. 5

Next are the performance issues. From an energy-consumption perspective, computation is very cheap: the power consumed by matrix floating-point operations is very cheap relative to data movement. Against the backdrop of widespread energy scarcity in North America, the performance bottleneck of AI systems has, in essence, shifted to the memory wall and to high-performance network interconnect and storage access. In fact, if you look at the microarchitecture of OpenAI's Jalapeño ("Redesigning the Inference Chip: From the Flaws of Nvidia GPUs to OpenAI Jalapeño"[3]), it too is essentially addressing the cost of memory movement: a unified L1 eliminates the complex data movement among TMEM / SMEM / RF, and avoiding a global L2 likewise reduces the cost of data movement.

Therefore, improving performance essentially comes down to solving the memory wall and high-performance network interconnect and storage access.

2.1 AI Agent Infrastructure

In agent scenarios, we have deeply optimized for agents' elasticity and security requirements. At this year's Apsara Conference, we released the Agent Sandbox[4] to build the foundational substrate for agent infrastructure. Through the collaborative optimization of the CIPU Cloud Infrastructure Processing Unit and virtualization technology, it provides each Agent with a MicroVM-level secure isolation environment: sandbox creation and cold start have a P99 of less than 180 milliseconds; after the resource pool is warmed up, warm start has a P99 of less than 20 milliseconds; it supports pause, deep hibernation, and fast wake-up, with a deep-hibernation wake-up time of less than 600 milliseconds.

In addition, the Agent Sandbox supports creating 100,000 sandboxes per minute, meeting the needs of model reinforcement learning (Agentic RL) and model RSI (recursive self-improvement). It provides millisecond-level delivery, highly concurrent startup, and burst-style elastic scaling for millions of agents, enabling models to learn faster and more efficiently and continuously improving model intelligence.

Figure 6
Fig. 6

Thanks to CIPU's elastic bare-metal technology, both Alibaba Cloud's own Agent Sandbox and DeepSeek Elastic Compute (DSec) can be fully accommodated. We note that DeepSeek also released a paper on elastic compute today, "DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale"[5], which is essentially also the logic of pooling multiple compute instance types together:

Figure 7
Fig. 7

It is worth noting in particular that we may be the only CSP in the world able to let DeepSeek's entire elastic-compute infrastructure migrate to the cloud directly without changing a single line of code. Comprehensive RDMA capabilities and local-disk virtualization instances allow 3FS to migrate to the cloud, while all CPU instances from the 8th generation onward support eRDMA and are fully compatible with the RDMA RC Verbs ecosystem, enabling foundation-model vendors like DeepSeek to keep the same software stack while elastically scaling out on the cloud.

2.2 GPU Bare-Metal Multi-Tenant Secure Isolation

SemiAnalysis's article discusses the security problems of Neo Clouds across multiple dimensions, such as firmware security for BIOS/BMC, BMC secure isolation, secure erasure of server firmware and memory/disk after a user terminates their lease, DPU/CIPU security isolation domains, DPU/CIPU software-upgrade isolation, and TPM firmware trust. This effectively exposes the product flaws of a certain DPU company—a company that will only begin to understand, in its next-generation DPU, why the host must be placed inside an untrusted domain—while the current generation of GPU bare-metal instances, built at large scale, is exposed to a great many security risks.

Figure 8
Fig. 8

CIPU, however, fully solved these problems back when it released its first-generation bare-metal instances in 2017.

2.3 AI Inference Infrastructure

In inference scenarios, CIPU supports multi-tenant high-performance access to multiple storage types, such as the AI inference KV Cache and model data loading, and, based on eRDMA, offloads functions like engram/PLE via CPU instances to reduce memory overhead on GPU servers. In the storage domain, through the Virtual Storage Channel (VSC) we have substantially improved the high-performance access and multi-tenant secure isolation capabilities of cloud storage such as CPFS/DFS, and, based on the standard NVMe KV command-set interface, we provide a native KV-interface-based KVCacheStore capability. We have also provided eRDMA data pass-through capability for OSS, becoming the world's first operator whose overlay RDMA traverses the cloud LB, with multi-tenant VIP service-based access to underlay cloud services.

Figure 9
Fig. 9

Beyond this, there is a great deal of RDMA data-transfer optimization involving PD (prefill-decode) disaggregation, and even cross-AZ transfer capability—that is, what Nvidia calls Scale-Across—yet we can deliver high-performance service without relying on any special switching fabric.

Figure 10
Fig. 10

2.4 AI Training Infrastructure

For training scenarios, what the entire industry is really doing is implementing lossy + multi-path transport to solve the reliability problems and load-imbalance problems of traditional RoCE networks. To achieve multi-path, AWS implemented the Reliable Datagram interface, which imposed enormous adaptation costs on the entire ecosystem. And even the MRC jointly built by OpenAI and Nvidia supports only a subset of RDMA semantics. This is in fact a very difficult problem to solve: multi-path forwarding causes out-of-order arrival at the receiver, and when the network becomes congested, one faces the dilemma of whether to switch to another path or to slow down the current path; it also requires extra probes to work with the switches to detect congestion. CIPU eRDMA solves this problem with a very simple algorithm, imposes no requirements whatsoever on the switches in the network, adapts well to networks with different oversubscription ratios, and can also fully utilize bandwidth in cross-AZ scenarios. In implementing these capabilities, it also ensures full compatibility with the standard RDMA RC Verbs interface, so that today's common communication ecosystems such as NCCL / IBGDA / DeepEP can be easily migrated to the cloud.

Figure 11
Fig. 11

Another problem is the reliability of traditional RoCE: reliability issues such as PFC storms brought on by lossless operation and transient fiber-optic link failures greatly affect the stability of the entire training cluster. During training, network-reliability-induced interruptions such as NCCL Hang frequently occur; therefore, the industry has gradually migrated to the lossy mode and chosen to use SACK in place of the original Go-back-N mechanism to better recover lost packets.

The principle is simple, but doing it well is extremely hard. For example, a well-known vendor shared at this year's Hot Chips that, under 1% packet loss, a 400 Gbps NIC can achieve an actual transfer performance of 160 Gbps with the help of SACK—what they consider to be state-of-the-art (SOTA) performance. What we would like to say is that three years ago, when CIPU 2.0 was completed, we had already achieved 90% goodput under 5% packet loss, reaching an industry-leading level.

Finally, returning once more to the cloud business perspective: we need to use the cloud's virtualized network (VPC) to achieve multi-tenant secure isolation + flexible networking, with no dependence on physical network topology. Once these capabilities are realized, they bring enormous benefits to customers:

Figure 12
Fig. 12

If interested, you may refer to "A Casual Chat on RDMA Modernization" to learn more..

Globally, the only one able to fully realize all four of these capabilities is Alibaba Cloud's CIPU—and this is still a chip designed three years ago. Somewhat regrettably, the industry has yet to catch up after more than three years.... So what even more advanced technologies will the next-generation CIPU bring? There's an Easter egg....

Figure 13
Fig. 13
References
[1]
An In-Depth Technical Analysis of the Cloud Infrastructure Processing Unit (CIPU): https://yunqi.aliyun.com/2026/session?agendaId=164
[3]
Redesigning the Inference Chip: From the Flaws of Nvidia GPUs to OpenAI Jalapeño: https://zartbot.github.io/blog/arch/jalapeno/index.html
[4]
Alibaba Cloud Agent Sandbox Officially Launched, Building the Foundational Substrate for Agent Infrastructure: https://mp.weixin.qq.com/s/oJwamxNmU3elJbQFe1dVAw
[5]
DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale: https://arxiv.org/abs/2609.22978