"After GPU comes memory, then power and cooling"... The step-by-step flow of AI bottlenecks
As the AI industry grows, the primary bottleneck is shifting from GPUs to memory, then to power supply, and finally to cooling systems.
Amidst the rapid growth of the AI industry, the key point for investors to focus on is the movement of the 'bottleneck' phenomenon. While the GPU (Graphics Processing Unit) was the core bottleneck of early AI infrastructure, as technology evolves, that bottleneck is sequentially moving to memory, power, and cooling systems.
The 'Bottleneck Investment Method' that distinguishes technology layers
According to a Tech Times TV video, tech writer Kim Ji-hyun (Vice President of SK SuperX AI Committee) emphasizes that when investing in technology companies, one must analyze them by dividing them into 'layers.' There are core technologies which are the source technologies, companies that commercialize them, and companies that supply solutions or provide services for those products. Using the era of the web as an example, if computer manufacturers and CPU suppliers generated profits in the early stages, service companies like Google or Naver made money later, and it has now evolved into a stage where companies like Facebook or TikTok generate revenue through advertising models.
Kim explained, "Since not all technologies create all innovations, timing is important," adding, "Rather than short-term trading aiming for quick profits, one must invest with a long breath of at least one to three years based on belief in the technology and actual performance." She mentioned, "I invested so well that I earned more from investments than from the company itself," and added that analyzing not only technology companies but also consumer goods or traditional industry companies that create new value using that technology is also one of the important strategies.
The flow from GPU to memory, and then to NAND flash
In a situation where computing resources are essential to improve the quality of AI models, the first bottleneck was the GPU. However, as the performance of GPUs increases, the importance of the memory that must fill the space next to them is growing, and the second bottleneck is moving to memory. This is because the number of HBM (High Bandwidth Memory) units mounted per GPU is increasing from 4 in the past to 8 currently, and the technology is evolving in a way where memory is stacked vertically rather than being attached to the side.
The phenomenon occurring in this process is the improvement in profitability for memory companies. As demand for HBM surges and memory manufacturers focus on HBM production, the supply of general DRAM relatively decreases. This is analyzed to lead to a price increase for DDR (Double Data Rate) DRAM required by smartphone and PC manufacturers such as Apple. Kim explained the background of how the profit margins of the three major memory companies have improved from less than 10% in the past to over 70% currently, stating, "If the supply of DRAM decreases due to HBM, DDR prices have no choice but to rise." Additionally, the spread of 'on-device AI' devices, such as the upcoming iPhone 18 Pro, is expected to stimulate demand for DRAM once again along with NPUs (Neural Processing Units).
Memory demand occurs not only for computation but also for storage. Since RAM is volatile memory where data disappears when the power is turned off, NAND flash and SSDs, which are non-volatile memory, are essential for storing vast amounts of data. Kim stated, "As data increases, if the desk (RAM) becomes insufficient, you eventually have to store it in a bookshelf (NAND flash)," and noted that the prices of storage devices made by Samsung Electronics and SK hynix are also showing an upward trend following the increase in demand.
Power supply and cooling systems: The emergence of the next bottleneck sections
It is predicted that after the memory issue is resolved, the bottlenecks will be power and cooling. This is because as the scale of data centers grows to the gigawatt (GW) unit, a level of power is required that is difficult to handle with existing power grids. Currently, the scale of South Korea's data centers is estimated to be around 1.5GW, but there are forecasts that it could increase to 20GW within the next 10 years. Since one GW consumes power equivalent to one nuclear power plant, securing new power sources such as green energy, Small Modular Reactors (SMR), and fuel cells is emerging as a key task rather than relying on the existing grid.
After the power supply is resolved, 'cooling' to dissipate the generated heat is expected to be the next bottleneck. As the computing speed of GPUs and HBM increases, the amount of heat generated becomes extreme, and existing air cooling (methods using fans, etc.) has the disadvantage of consuming too much power. Therefore, water cooling, which uses liquid to cool heat, or immersion cooling, which submerges the equipment entirely in liquid, are emerging as alternatives. Kim predicted, "The revenue of companies related to cooling solution parts, materials, and equipment can increase with a breath of about two years." This is because the construction of large-scale power infrastructure and data centers does not happen in a short period.
0Comments
Comments are currently disabled.