It's Not Computing Speed, but 'Data Movement' that Determines AI Performance... Paradigm Shift Toward Memory-Centric Infrastructure
The standard for performance competition in the AI era is shifting from simple computing speed to the efficiency of data movement. As AI workloads become more advanced, minimizing data movement through memory-centric architectures is becoming the key to system efficiency and cost reduction.
In the era of Artificial Intelligence (AI), the standard for performance competition is shifting from simple computing speed to 'the efficiency of data movement.' This is because how minimally and efficiently vast amounts of data can be moved has emerged as a key factor determining the overall performance and cost of AI systems.
'Bottlenecks' where data movement costs more than computation
Current computing systems are structured around a processor that fetches and processes data. The process involves data moving through storage, memory, and cache to the computing unit and then returning repeatedly. However, as AI workloads become more sophisticated, this processor-centric structure is facing limitations. As the scale of data and the frequency of access surge, the latency, power consumption, and hardware costs occurring during the data movement process have become decisive factors hindering system efficiency.
In actual modern computing environments, much higher costs are incurred in memory access and data movement than in the computation itself. The energy required to read data once from DRAM can be at least 150 times to as much as 2,000 times greater than simple arithmetic operations. Especially in the case of Large Language Models (LLM), due to the characteristic of needing to refer to previous context every time a token is generated, the demand for memory capacity and bandwidth rises sharply, further intensifying memory bottlenecks. Accordingly, the focus of the AI infrastructure competition is moving beyond securing faster computing units toward co-designing computing units, memory, storage, and networks as a single system to minimize data movement.
CXL and NVLink, the bridge connecting memory expansion and high-speed communication
The core technology that has emerged to solve data movement bottlenecks is the high-speed Interconnect. This serves to widen the pathways between devices and flexibly connect memory and computing units according to the workload. Representative technologies include CXL (Compute Express Link) and NVLink/NVSwitch.
CXL expands the connections between processors, accelerators, and memory devices to flexibly increase memory capacity and provides a foundation for 'memory pooling' where multiple devices can share memory. On the other hand, NVLink and NVSwitch are responsible for high-bandwidth connections within GPU clusters and focus on reducing the data bottlenecks that occur when large-scale AI models move between multiple GPUs. Although these two technologies have different application areas, they share the common goal of improving the data movement process to utilize system resources efficiently. Once this technological foundation is established, memory and accelerators will function as system resources that can be freely combined according to the workload, rather than components dependent on specific devices.
The rise of 'Near-Memory Accelerators' that process data directly near memory
To overcome the limitations of merely increasing connection speeds, 'near-memory acceleration' methods, which process data directly near the memory instead of bringing it to the computing unit, are gaining attention. This aims for an 'interconnected near-memory accelerator architecture' where, instead of the central computing unit (GPU/NPU/TPU) processing all data, accelerators around the memory perform assigned tasks and exchange only the necessary results.
This method is effective in reducing latency and power consumption by reducing data round-trips in 3D stacked memory environments such as HBM (High Bandwidth Memory). According to the 'Tesseract' design, a related research case, when multiple near-memory accelerators were connected via high-speed interconnects and applied to parallel graph processing, it showed approximately 10 times the performance improvement and a similar level of energy saving compared to existing processor-centric designs. Recently, various attempts to innovatively reduce data movement are continuing, such as the CENT design research that combines memory expansion structures with near-memory processing targeting LLM inference.
0Comments
Comments are currently disabled.