Edge AI Moves Closer—and Memory Matters More Than Ever
Today, AI is rapidly expanding beyond large-scale data centers into smartphones, AI PCs, vehicles, robots, and industrial environments. As cloud AI and edge AI divide roles based on their respective strengths and grow together, a new AI infrastructure cycle is beginning.
Amid this transformation, memory is evolving beyond a component that simply stores data into a core technology that determines the performance and energy efficiency of AI systems.
The Expansion of the Cloud Ecosystem and a New AI Infrastructure Cycle
With the emergence of diverse services based on generative AI and inference AI, artificial intelligence is quickly becoming part of everyday life. Until recently, AI infrastructure investment was concentrated primarily in large-scale cloud data centers operated by hyperscalers. AI is now expanding in earnest into small and medium-sized businesses, industrial sites, and individual user environments.
High-performance AI workstations, on-premises AI servers installed directly at corporate and industrial sites, AI PCs, and smartphones are representative examples. The emergence of desktop-sized personal AI supercomputers also reflects this trend.
At the center of this movement, which brings AI closer to where data is generated, is edge AI. Rather than sending data to a centralized cloud for processing, edge AI performs all or part of the AI computation on devices such as smartphones and PCs or on local systems installed in vehicles, factories, and retail environments. On-device AI, which runs AI models directly on the device, is one of the most representative forms of edge AI.
Real-time translation and image editing on smartphones, personal assistant capabilities on AI PCs, environmental perception in vehicles, and equipment anomaly detection in smart factories already demonstrate how edge AI is being applied across everyday life and industrial environments.
This does not mean that edge AI will replace cloud AI. The cloud will continue to handle workloads requiring large-scale model training and high scalability, while the edge will increasingly process tasks that require rapid response, data control, and immediate action in the field. Hybrid architectures, in which the cloud and edge divide and process a single workload according to their respective strengths, are also continuing to evolve.
Ultimately, AI infrastructure is not simply shifting from the cloud to the edge. Instead, it is evolving in a direction where cloud and edge environments expand in a complementary manner.

Three Core Values of Edge AI
The defining characteristic of edge AI is that intelligence is located closer to users and the data they generate. This creates three primary benefits.
① Ultra-Low-Latency, Real-Time Processing
Edge AI can reduce response latency by minimizing the need to send data to a remote server and wait for the results to return. This is particularly important for applications that require immediate decisions and responses, including hazard detection in autonomous vehicles, real-time translation, immersive gaming, robot control, and anomaly detection in industrial equipment.
② Data Security and Privacy
Security-sensitive data—such as personal voice recordings and photographs, manufacturing data, and confidential corporate information—can be processed within the device or local system without being transmitted to an external server. This allows users to access AI services while managing their data more securely, while businesses can apply AI to their operations without relinquishing control over their information.
The actual level of security depends on the overall hardware, software, network, and data-management framework. Nevertheless, local processing offers a clear advantage by minimizing the external transmission of sensitive source data.
③ Reduced Cloud Costs and Network Burden
Sending all data to centralized servers can increase network traffic and cloud infrastructure operating costs. By processing repetitive or time-sensitive tasks at the edge, systems can reduce both the cloud’s computational workload and the volume of data transmitted over the network.
Even in environments with limited network connectivity, devices can continue to provide essential AI functions by processing data locally without relying on the central cloud. This can improve both system reliability and the range of environments in which AI can be deployed.
Edge AI Expands Toward Agentic AI
Edge AI is moving beyond simply responding to user commands. It is evolving toward agentic AI, which can understand user intent and surrounding conditions, then independently plan and execute multiple tasks.
In an agentic AI environment, a single AI model does not operate alone. Multiple agents may simultaneously analyze information, select tools, and coordinate the sequence of tasks. As a result, the role of the CPU—which manages the overall system and multiple workloads—becomes increasingly important alongside AI accelerators.
However, increasing the computing performance of xPUs such as CPUs, GPUs, and NPUs does not automatically deliver a proportional improvement in overall AI system performance. No matter how fast the processor may be, the xPU must wait if memory cannot supply the required data at the right time.
The actual performance of an AI system depends not only on computing power, but also on the capacity to store data, the bandwidth to transfer it, and the energy efficiency required to achieve both within a limited power budget.
Edge AI systems, in particular, operate in environments with more limited space, power, and cooling resources than large-scale data centers. The high bandwidth, high capacity, and energy efficiency traditionally required by data center-class AI systems are now becoming necessary in relatively compact devices such as AI PCs, workstations, vehicles, and robots.
Three Challenges Facing the Edge AI Era
① Capacity for Larger AI Models
As AI models become more advanced, the number of parameters that make up those models and the amount of data generated during execution also increase. To deliver high-performance AI in a local environment without continuous reliance on the cloud, memory must be able to accommodate more model weights and contextual data.
When memory capacity is insufficient, the system must repeatedly transfer the required data between storage and memory or load only selected parts of a model at a time. This process can increase response times and consume additional power.
High-capacity memory is therefore a fundamental requirement for enabling edge AI systems to run larger and more sophisticated models reliably in local environments.
② Bandwidth to Fully Utilize xPU Performance
Bandwidth indicates how quickly memory can supply data to an xPU. Generative AI and agentic AI continuously read and process large volumes of model weights and contextual data while interacting with users.
When memory bandwidth is insufficient, it becomes difficult to fully utilize the xPU’s computing resources. Delivering fast and seamless AI responses requires high memory bandwidth suited to the workload, as well as an efficient data-movement architecture.
The required bandwidth varies depending on model architecture, precision, batch size, and target response speed. What is clear, however, is that AI system bottlenecks are expanding beyond computing performance to include data movement between the xPU and memory.
③ Efficiency Within Limited Power and Space
Unlike data centers, edge AI systems operate with limited power and cooling resources. Smartphones and AI PCs must account for battery life, while vehicles and robots must perform multiple functions within a fixed power budget. Small-scale AI servers installed in the field must also consider installation space, heat generation, and cooling costs.
In particular, repeatedly moving data among storage, memory, and processing units consumes a significant amount of power. AI systems must therefore be designed not only to increase processing speed, but also to reduce the distance and frequency of data movement and process more data per watt.
High energy efficiency is essential not only for extending the operating time of battery-powered devices, but also for reducing heat generation, cooling requirements, and overall system operating costs.

New Memory Solutions for AI at the Edge
For edge AI to execute larger and more sophisticated models quickly and efficiently, capacity, bandwidth, and energy efficiency must all improve together. This is a complex challenge that cannot be addressed simply by enhancing the individual performance of conventional memory products.
New memory architectures are therefore required—architectures that shorten the physical distance between processors and memory, minimize data movement itself, and connect DRAM, NAND, logic, and packaging from a system-level perspective.
In Part 2, “Samsung Memory Solutions Connecting the Future of Edge AI,” we will explore how next-generation solutions such as zHBM, LPDDR-PIM, and zNAND-O, together with Samsung’s integrated capabilities, can address these challenges.
* The contents of this page are provided for informational purposes only. No representation or warranty (whether express or implied) is made by Samsung or any of its officers, advisers, agents, or employees as to the accuracy, reasonableness or completeness of the information, statements, opinions, or matters contained in this page, and they are provided on an “AS-IS” basis. Samsung will not be responsible for any damages arising out of the use of, or otherwise relating to, the contents of this page. Nothing in this page grants you any license or rights in or to information, materials, or contents provided in this document, or any other intellectual property.
Custom range