← Back to list

CXL Unpacked: Enabling Scalable, Flexible, and Fast Memory for the AI Era

In our previous article, “CXL in Action: Overcoming Traditional Scale-Out Limits with MX1,” we introduced the emergence of Compute Express…

Jay in XCENA BLOG · 2025-08-08 01:36 · 3 claps · 5.1 min read
#cxl #ai #accelerator #data-center #processing-near-memory
Open on Medium ↗
Wiki topics: AI · AI · General STP · Startups & Venture 🌐 · Web Development 📰 · Journalism & News

CXL Unpacked: Enabling Scalable, Flexible, and Fast Memory for the AI Era

In our previous article, CXL in Action: Overcoming Traditional Scale-Out Limits with MX1,” we introduced the emergence of Compute Express Link (CXL) and how it is set to transform memory infrastructure. In this follow-up, we take a closer look at how CXL works, explaining its core protocols (.io, .cache, and .mem), the different device types (Type 1–3), and the major features it enables, such as memory expansion, pooling, and sharing. We also examine some of CXL’s current limitations and discuss emerging strategies to address them.

From PCIe to CXL: The Evolution of Data Communication

CXL builds upon the proven foundation of PCIe, a serial communication interface widely adopted in data centers. PCIe offers several critical advantages:

  • High-Speed Connectivity: PCIe delivers substantial bandwidth improvements, enabling rapid data transmission with minimal latency.

  • Simplified Infrastructure: Serial interfaces reduce complexity by using fewer physical connections, leading to enhanced reliability and lower costs.

  • Scalability and Efficiency: PCIe scales effectively, accommodating increased bandwidth and higher device densities without sacrificing signal integrity.

Unlike older parallel bus architectures that transmit multiple bits simultaneously across many physical pins, PCIe uses high-speed serial links to transmit data over fewer pins, resulting in higher bandwidth efficiency per pin.

By leveraging PCIe, CXL capitalizes on these strengths, introducing additional protocols designed specifically to optimize memory and computational performance.

CXL Protocols and Device Types

CXL-enabled devices communicate with the host system using a combination of three protocols: CXL.io, CXL.cache, and CXL.mem. These protocols allow devices to be configured and managed, access the host CPU’s memory efficiently, and provide additional memory capacity to the system. Specifically, CXL.io handles setup and communication tasks similar to traditional PCIe, ensuring smooth integration and control. CXL.cache enables devices like accelerators to coherently access and cache the host processor’s memory, reducing latency and improving performance. CXL.mem allows the host to directly access memory attached to the device, significantly expanding system memory capacity.

Based on which protocols they support, CXL devices fall into three categories. Type 1 devices, like network interface cards (NICs), use CXL.io and CXL.cache to access host memory but do not include memory of their own. Type 2 devices, such as GPUs and custom accelerators, support all three protocols, enabling them to cache host memory and also provide their own memory to the system. Type 3 devices are focused solely on memory expansion and support CXL.io and CXL.mem, supplying large pools of memory to the host without coherent caching.

Source: Trends in Compute Express Link(CXL) Technology (Seonyoung Kim et al., ETRI, 2023)

Source: Trends in Compute Express Link(CXL) Technology (Seonyoung Kim et al., ETRI, 2023)

Major Features of CXL: Cache Coherence, Memory Expansion, and Beyond

CXL supports robust cache coherence, a critical feature that ensures the host CPU and attached devices maintain a synchronized and up-to-date view of shared memory. In traditional systems, inconsistencies can arise when multiple components cache the same data independently. CXL solves this by enabling hardware-based coherence protocols, allowing accelerators such as GPUs, FPGAs, or custom AI processors to read from and write to the same memory regions as the CPU without introducing data mismatches or requiring complex software-level synchronization. This hardware coherence dramatically simplifies programming models and accelerates performance for workloads that require tight coordination between the CPU and accelerators.

Beyond coherence, CXL also addresses one of the most pressing limitations in server architecture: memory capacity scaling. Most servers are constrained by the number of memory slots on the motherboard and the capacity limitations of DIMMs. CXL breaks this barrier by enabling external memory modules to function as if they were local memory. This allows system builders to go well beyond traditional DRAM limits and deploy terabyte-scale memory configurations in compact, power-efficient form factors.

In addition to expanding capacity, CXL improves memory bandwidth, or the rate at which data can be transferred between memory and compute elements. By offloading memory access from the CPU’s native memory controller to high-speed CXL-connected devices, overall system throughput is improved. This is especially valuable for memory-bound applications that depend on rapid and sustained memory access rates.

Another game-changing feature of CXL is its support for switch-based memory expansion and sharing. Just like Ethernet switches allow multiple devices to share a network, CXL switches enable multiple servers and devices to share memory resources over a fabric. This makes memory pooling possible: memory can be dynamically allocated to whichever server needs it most, reducing idle capacity and avoiding overprovisioning. In hyperscale data centers, this creates a more flexible and efficient memory infrastructure. It also enables memory sharing, where multiple hosts can simultaneously access a common memory pool — for example, in distributed analytics workloads or AI training pipelines where large datasets must be shared in parallel.

Together, these features make CXL a foundational technology for next-generation computing environments — where flexibility, scale, and performance must go hand-in-hand. From AI-driven inference tasks to cloud-native applications and large-scale simulations, CXL provides the architectural tools to meet the growing demand for fast, elastic, and coherent memory systems.

Source: CXL Consortium

Source: CXL Consortium

Overcoming the Setback: Near-Data Processing (NDP) and XCENA’s MX1

While CXL offers transformative benefits for memory architecture, it also introduces challenges, particularly increased latency and system complexity due to external memory access and the need to manage cache coherence. To address these issues, Near-Data Processing (NDP) brings compute capabilities closer to the data itself. This approach reduces the need to move data back and forth between memory and the processor, significantly cutting down latency and power consumption. NDP is especially valuable for workloads like real-time analytics, machine learning, and AI inference, where performance is often limited by data movement rather than raw compute power.

XCENA’s MX1 is specifically engineered to mitigate latency concerns inherent to CXL-based memory expansion:

  • Integrated On-chip Cache: MX1 features a large on-chip cache, substantially reducing access latency by caching frequently accessed data closer to processing resources.

  • NDP Capabilities: MX1 includes 1,000s of specialized NDP cores designed to perform data-intensive computations directly near memory. Tasks such as vector search, data analytics, and AI workloads benefit significantly from these embedded processing units, dramatically cutting data movement and boosting overall system performance.

Source: XCENA Inc.

Source: XCENA Inc.

CXL represents a transformative evolution in memory architecture, leveraging the strengths of PCIe’s high-speed serial interface to unlock new levels of scalability, flexibility, and performance. By enabling coherent, composable, and pooled memory architectures, CXL is reshaping how data centers and compute systems are designed, paving the way for more agile, efficient, and AI-ready infrastructure.

XCENA’s MX1 stands at the forefront of this shift. By combining CXL-based memory expansion with intelligent features like integrated L3 cache and Near-Data Processing, MX1 offers a practical and forward-looking solution to the latency and complexity challenges that arise in modern data-intensive workloads.

To learn more about how XCENA is shaping the future of memory infrastructure, visit our website at xcena.com and follow us on LinkedIn for the latest updates, demos, and industry insights.


메타데이터
post_id
e3e9df3c08ea
slug
cxl-unpacked-enabling-scalable-flexible-and-fast-memory-for-the-ai-era-e3e9df3c08ea
url
https://medium.com/xcena-blog/cxl-unpacked-enabling-scalable-flexible-and-fast-memory-for-the-ai-era-e3e9df3c08ea
canonical_url
https://medium.com/xcena-blog/cxl-unpacked-enabling-scalable-flexible-and-fast-memory-for-the-ai-era-e3e9df3c08ea
author_url
https://medium.com/@jaewoo.yeon
status
ok
fetched_at
2026-06-11 05:11:55