When encryption, compression, indexing, memory, and context services are positioned close to the medium, the value boundaries of SSDs, DRAM, HBM, and HBF will be redefined.
The AI agent is transforming computers from systems that perform a single model invocation upon receiving a question into continuous operating systems that observe, reason, invoke tools, modify environments, and retain state.
A single task may involve sequential access to models, memory, vector indexes, business databases, object storage, and network services, while continuously writing back tool outputs, execution traces, user preferences, temporary context, and audit logs. The longer an agent runs, the more its value depends on the secure storage, accurate retrieval, timely updates, and low-cost reuse of data [1].
This means storage is no longer just the final destination for data after an Agent completes its task, but will gradually become integrated into the Agent’s perception, memory, and decision-making cycle. AI SSDs have already demonstrated the first directions: storing models and adapters, absorbing KV caches evicted from memory, reducing cold start times, and increasing the capacity of deployable models. Taking the next step, SSDs may also perform encryption, compression, deduplication, tagging, indexing, versioning, and lifecycle management during data persistence, while providing long-term capacity for Agent Memory.
This change did not emerge out of nowhere. Encryption and decryption of encrypted drives have long been performed transparently within controllers; the Compute Storage standard lists compression, encryption, regex filtering, and erasure coding as functions that can be executed at the storage side; products such as Samsung SmartSSD have previously offloaded database scanning and video processing to the drive level [4][5][7]. The novelty of the Agent era is that these capabilities are no longer solely dedicated to general-purpose data processing, but are now being reconfigured around agent identity, context, memory, tool trajectories, and Token costs.
In future markets, different terms such as "secure SSD," "compressed SSD," "retrieval SSD," "memory SSD," or "contextual SSD" may emerge, or they may not evolve into separate hardware categories at all, instead converging into programmable, functionally oriented SSDs: maintaining backward compatibility with standard storage at the foundational level, while upper-layer software discovers and invokes device capabilities based on specific use cases.
NVMe has established command sets such as Computational Programs and Subsystem Local Memory, providing a standardized pathway for device-side program discovery, configuration, and execution [6]. The real industry challenge is not merely whether to embed a processor within the drive, but who defines the data semantics, functional boundaries, and end-to-end outcomes.
The agent is not a single call.
but rather a continuously written data chain
The primary persistent objects of traditional chatbots are conversation logs, whereas agents generate more complex data graphs. They must store observations, tool outputs, plans and reflections, intermediate task states, user profiles, environment snapshots, retrieved evidence, and execution logs; on the model side, they also produce KV Cache, Prefix Cache, Adapters, expert weights, and checkpoints. These different objects vary in their update frequencies, reuse scopes, and security levels, yet collectively determine whether subsequent reasoning can be sustained.
Agent Memory does not equate to stuffing all historical conversations into a vector database. Recent systematic research from a data management perspective has broken down Agent Memory into four modules: representation and storage, information extraction, retrieval and routing, and maintenance, and has shown that no single architecture dominates across all workloads [9].
Systems like Mem0 also emphasize converting raw conversations into more compact, reusable long-term memories to reduce input token usage and retrieval burden in lengthy conversations [10]. This means that, in the future, the storage layer must preserve not only content but also how that content is organized, updated, and forgotten.
Therefore, Agent data persistence requires a richer contract than merely "write success." A memory object should include user or Agent identity, tenant, source, timestamp, version, permissions, trustworthiness, retention period, and deletability; model context must also be bound to the model version, tokenizer, positional encoding, and adapter. Only when namespace and lifecycle are explicitly defined can subsequent compression, indexing, caching, and sharing avoid violating semantic boundaries.
The appropriate role of a functional SSD within this data chain is not to independently determine whether an experience is worth remembering, but rather to perform deterministic tasks close to the data after the Runtime has already provided the object and strategy. For example: selecting keys by tenant, placing data according to lifecycle, compressing and deduplicating reducible objects, maintaining index pages and metadata, prioritizing prefetching of data likely to be used in inference, and feeding back tail latency, write amplification, and media health to the upper layers. Semantic decisions remain with the Agent and Runtime; data execution should occur as close to the medium as possible.

Figure 1: Capability spectrum of functional SSDs in the Agent era. Security, content reduction, indexing, memory, reasoning context, and governance functions can be embedded in the device or implemented collaboratively by downloadable programs, runtimes, and storage nodes. This figure illustrates industry mechanism projections, with boundaries for compute, storage, and security capabilities referencing SNIA and NVMe specifications [4][5][6].
Core insight: The SSD upgrade in the Agent era is not merely about increasing on-disk compute power, but about combining standard block devices, discoverable in-data functions, object semantics, and lifecycle governance. Functions closer to the medium should be deterministic, auditable, and isolatable; decisions closer to the model should remain in the runtime.
Encryption is already in place; the change is that the policy now follows the Agent.
"Automatic data encryption at rest" is not a futuristic concept. Self-encrypting drives use dedicated hardware within the controller to encrypt data upon writing and decrypt it upon reading, enabling transparent static data protection through key management and policy controls [5].
For Agent systems, the new requirement is to further refine encryption granularity from entire disks or single namespaces down to users, Agents, tasks, and objects: personal memories and application caches on the same edge device must not exceed their respective permissions, and shared contexts among multi-tenant Agents on the cloud must clearly define which data can be reused and which can only be accessed within a single permission domain.
The NVMe specification has introduced finer-grained capabilities such as host-managed keys and Key Per I/O [6], providing the foundation for carrying different security contexts with each I/O operation. Future secure SSDs may combine data provenance, timestamps, access logs, integrity verification, and trusted deletion, enabling agents not only to answer “What do I remember?” but also “Where did this memory come from, has it been modified, who accessed it, and when must it be deleted?” For finance, healthcare, enterprise knowledge bases, and personal AI, the chain of evidence may be as critical as retrieval speed.
Encryption also alters the order of other near-data functions. Encrypted byte streams are typically difficult to compress and deduplicate efficiently, so content reduction should generally precede encryption [5]; indexing requires clear boundaries between plaintext characteristics, protected metadata, and searchable ciphertext. The competitiveness of functionally enhanced SSDs lies not merely in the number of features they offer, but in their ability to correctly orchestrate compression, indexing, encryption, storage, and deletion through a verifiable pipeline while avoiding any step that expands the attack surface.
Compression, indexing, and memory maintenance may become the next set of SSD features.
Compression is one of the easiest features to form a commercial closed loop. Agents repeatedly write text, JSON, logs, vectors, checkpoints, and multimedia intermediate results, a significant portion of which exhibit structural redundancy. If compression is performed before data enters the network or NAND, it can reduce transmission, physical writes, and storage consumption, indirectly lowering energy usage and media wear. However, compression ratio, additional latency, CPU savings, and write amplification must be evaluated together; further compressing already quantized or highly compressed model weights may yield limited benefits.
The indexing function has a more direct relationship with agents, as memory is only valuable when correctly retrieved. KIOXIA AiSAQ places vectors and indexing structures on SSDs, reducing DRAM usage through SSD-friendly clustering and graph search, and has demonstrated billion-scale vector retrieval on a single server [8]. It is essential to clearly distinguish that AiSAQ is primarily a software technology using SSDs as the main indexing medium—it does not imply that standard SSDs automatically generate embeddings or understand semantics. A more likely industry pathway is for GPUs, NPUs, or CPUs to generate representations, while SSDs and near-data processing handle indexing organization, candidate filtering, and returning smaller result sets.
Memory maintenance is more complex than indexing. Long-term agent memory involves additions, mergers, conflicts, revisions, de-prioritization, expiration, and forgetting; the same fact may simultaneously exist in multiple forms—original records, summaries, vector representations, and knowledge graph embeddings [9]. Future "memory SSDs" could provide atomic updates, logging, TTL, cold-hot placement, and secure deletion for these versions, but similarity alone cannot determine what constitutes a true memory. Memory quality still depends on upper-layer processes for extraction, routing, conflict resolution, and evaluation.
From a product perspective, these functions do not necessarily require complex models to run on each drive. Compression, encryption, hashing, filtering, index page maintenance, and object lifecycle management are well-suited for deterministic dedicated circuits or lightweight programs; whereas embedding, reordering, summarization, and memory merging may be handled by the host or standalone accelerators.
The key to a functional SSD is connecting the two types of work using a unified object ID and an observable interface, thereby reducing data movement—rather than cramming all AI computations into the drive.
On-device: SSDs may become the long-term state layer for personal agents.
On-device agents will continuously access personal documents, photos, emails, calendars, app states, browsing history, and device sensor data. Unified memory is suitable only for retaining the current working set, while SSDs can store larger local model libraries, adapters, vector indexes, personal memories, and tool trajectories. As long as they maintain the standard NVMe form factor, functional SSDs can first be installed directly as standard system drives in AI PCs and workstations, with security, indexing, and context capabilities gradually enabled through drivers, runtimes, and firmware.
Its value to consumers isn't that "the hard drive thinks," but rather that local AI can remember longer, continue operating under weak network conditions, and reduce the amount of raw data and input tokens uploaded to the cloud for each task. A meeting agent can retain audio, summaries, and task status; a programming agent can maintain repository indexes and modification histories; a home agent can share authorized photos, documents, and device statuses across devices. SSDs keep these states nearby, and the router decides whether to process locally or invoke the cloud based on quality, privacy, power consumption, and network conditions.
The edge market may also expand from a single drive to small storage nodes. AI PCs, home servers, or store-edge boxes can provide local model replicas, personal memories, vector databases, and encrypted archives for smartphones, tablets, robots, and cameras, avoiding redundant storage of the same data across devices. Business models will extend beyond capacity upgrades to include AI PC premiums, local agent subscriptions, model and skill package management, and private AI nodes for homes and small businesses.
Local storage does not automatically equal privacy. If an application can freely read memory, old versions cannot be deleted from indexes, and keys are decoupled from device identity, more features only mean a larger attack surface. Endpoint functional SSDs must treat application isolation, user consent, retention periods, verifiable deletion, and key revocation after device loss as core product capabilities—not merely as marketing buzzwords for privacy.
Cloud side: SSDs will move from devices to Agent storage nodes.
The state of cloud-side agents is larger and requires greater sharing. A complex request involves chaining multiple model calls, tool executions, memory accesses, and network transmissions; multi-agent collaboration also generates shared plans, messages, evidence, and execution logs [1]. Local node SSDs can store models, checkpoints, and high-frequency indexes; rack- or POD-level Flash layers can accommodate shared KV, prefixes, adapters, and shared memory across GPUs; general-purpose object storage continues to hold cold data and long-term factual sources.
Mooncake has organized CPU, DRAM, SSD, and RDMA/NIC into a distributed KV Cache pool, enabling the scheduler to determine request paths based on cache location and TTFT/TBT targets [3]. NVIDIA CMX has established a POD-level Flash context layer for KV Cache across HBM, host memory, and general-purpose shared storage, integrating shared storage nodes into the inference data path [2]. The end-to-end benefits of these systems cannot be attributed to a single SSD, but they demonstrate that "storage nodes participating in token generation" is evolving from concept to infrastructure product.
The next-generation agent storage node may simultaneously provide context, memory, and governance services. It can maintain shared prefixes and KV directories, warm up models and adapters, store vector and graph indexes, compress and encrypt data by tenant, log tool calls and evidence chains, and expose metrics such as hit rate, P99 latency, write amplification, energy consumption, and media health to the scheduler. Customers will be purchasing not just disk capacity and bandwidth, but also GPU utilization, SLO goodput, tokens per second, tokens per watt, queries per second, and audit response speed.
This will also transform the business model. Single drives can still be sold based on capacity, durability, and performance; storage nodes may generate composite revenue through multi-drive appliances, control planes, runtime licenses, long-term support, and SLAs; at a higher level, services could emerge that bill based on contextual effective capacity, number of memory objects, retrieval throughput, or effective tokens. Once functional SSDs assume responsibility for end-to-end outcomes, their value proposition expands from semiconductor components to data infrastructure.

Figure 2: Typical division of storage responsibilities between edge-side and cloud-side agents. The edge side emphasizes personal data, local models, offline capabilities, and privacy control; the cloud side emphasizes shared context, multi-tenant governance, GPU utilization, and token economics. This figure represents a product形态 inference, with the cluster side referencing Mooncake and CMX[2][3].
The value boundaries of HBM, HBF, DRAM, and SSD will be redefined.
The storage rearchitecture brought by Agent should not be interpreted as SSD replacing memory. SK hynix’s Tiered Memory concept introduced at FMS 2026 connects HBM, DRAM, NAND/HBF, and SSD based on speed, capacity, and cost, with the goal of minimizing data movement [11]. This directly aligns with the reality of Agent systems: data currently in use, data to be used within the next few milliseconds, data that may be reused across sessions, and long-term archived data each require different storage media to efficiently handle them.
HBM remains the closest layer to GPU compute, holding current weights, activations, hot KV caches, and operator intermediate states, with metrics including tokens per second, bandwidth, and compute utilization. Agents, which enable longer context and multi-model collaboration, will further increase HBM demand rather than diminish its value. The constraints on HBM are capacity cost and supply; therefore, systems must move data that does not require immediate access to lower-cost layers and accurately prefetch it before use.
HBF, or High Bandwidth Flash, is a new tier that has rapidly emerged since 2025. In August 2026, SK hynix and SanDisk unveiled their first open specifications, covering capacities up to 512 GB, with bandwidth divided into three tiers ranging from approximately 0.4 TB/s to 3.0 TB/s, and utilizing UCIe to connect to processors [12]. HBF achieves larger capacities than HBM by using NAND flash, aiming to bridge the gap between HBM and SSD for read-intensive large inference workloads. It remains in the early stages of standardization and productization; official specifications do not equate to widespread production performance and cannot eliminate NAND’s inherent limitations in latency, write endurance, and variable-state access.
The H3 study proposes a clearer division of labor: place read-only data in HBF and retain other data in HBM to build a hybrid inference system leveraging the strengths of both [13]. For agents, HBF is better suited for model weights, MoE experts, and large working sets dominated by reads; frequently updated KV caches, activations, and runtime states remain better suited for HBM or DRAM. If HBF matures, it may reduce the need to increase GPU/HBM capacity, but it is more likely to serve as a complement rather than a replacement.
DRAM and CXL occupy an intermediate ground between variable states and shared capacity. Host KV, index hot sets, prefetch buffers, and agent states require low latency and frequent modifications; CXL expansion, pooling, and sharing can reduce memory silos and provide more flexible capacity for multiple hosts [14]. SK hynix demonstrated CXL pooled memory and DRAM-SSD hybrid solutions for KV sharing, prediction, and prefetching at FMS 2026, but these are specific demonstration results that require validation across more platforms [11].
Functional SSDs handle larger, more persistent, and more governance-intensive cold and warm objects: models, adapters, KV/prefixes, vector indexes, agent memory, logs, and checkpoints. Compared to HBF, SSDs are farther from computation but offer standardized form factors, mature ecosystems, and advantages in capacity and cost, making them better suited for node-level sharing. Future competition will not focus solely on media bandwidth, but on who can deliver the right objects at the right time while being accountable for effective tokens, retrieval throughput, and data governance.

Figure 3: Reallocation of roles among HBM, HBF, DRAM/CXL, functional SSDs, and shared storage in the Agent era. HBF specifications are based on the first open standards announced by SK hynix and SanDisk in August 2026 [12]; hierarchical logic references FMS 2026 Tiered Memory, H3, and CXL data [11][13][14].
AI SSD represents the first industrial samples of functional SSDs.
Current AI SSDs are largely trending toward two ends. One end consists of enterprise SSDs optimized for AI workloads, providing a foundational medium for model caching, checkpoints, and vector indexing through low latency, high IOPS, sustained bandwidth, durability, and capacity density. Representatives of this category include Innogrit’s Dongting N3X and Huawei’s OceanDisk LC 560 [19][20]. These primarily address whether AI data can be stably stored, continuously written, and promptly retrieved—they do not inherently possess Agent Memory, indexing, or contextual semantics simply because they target AI. The other end is beginning to actively participate in the inference data path: SSDs no longer merely receive generic block requests from the operating system, but instead, through middleware, runtime, or on-storage processing capabilities, gradually identify AI-specific objects such as model weights, experts, KV caches, and prefetch windows. Phison, Longsys, and Yinpu—Linkmicro are three representative R&D approaches in this direction.
Phison aiDAPTIV adopts a solution based on “mature SSD capabilities + middleware + complete toolchain.” aiDAPTIVLink manages memory between GPU VRAM and Flash, slicing model weights at runtime to prioritize active weights in VRAM while offloading inactive weights to aiDAPTIVCache SSD, and preserving evicted KV Cache to avoid recomputation after context eviction [15]. The commercial value of this approach lies in delivering an integrated solution combining SSDs, software licensing, deployment tools, and system support—enabling customers to expand model and context capacity on workstations or local servers without redesigning their compute chips. Its technical focus is on compatibility with existing GPU and AI software ecosystems, leveraging mature controllers, firmware, endurance design, and platform partnerships to reliably extend Flash as an additional capacity layer beyond VRAM.
Jiangbo Long’s SPU + iSA more closely resembles a combination of a "storage execution layer + software decision layer." The SPU handles lossless compression, HLC advanced caching, and data scheduling across different NAND media at the device level; the iSA makes decisions around MoE expert offloading, KV Cache lifecycle management, and Device Smart Prefetch, translating AI workload characteristics into prefetching, backfilling, compression, and hot-cold migration actions [16]. Thus, its focus is not merely on storing more data on the SSD, but on reducing DRAM usage, increasing effective capacity, and aligning storage-side operations with endpoint inference rhythms. This approach demonstrates that when greater execution capability is available near the controller, an SSD can evolve beyond being a simple block device into a near-data processing node capable of handling compression, caching, and object scheduling.
Yinpu adopts a co-defined approach between computation and storage. Yinpu analyzes expert activations, KV lifecycle, data precision, and computation windows, starting from model architecture, runtime, operating system, CPU/GPU/DRAM co-design, and whole-system reference designs. Lianyun maps these requirements onto SSD controllers, firmware, cache partitioning, queues, NAND adaptation, and mass production systems [17][18]. This enables product definition to be derived retroactively from token latency, model capacity, and end-to-end SLOs—determining when data should enter the SSD, at what granularity it should be stored, how to prefetch it, and which operations are better suited to be offloaded closer to the controller, rather than making only localized optimizations on existing block I/O. Compared to the first two approaches, this cross-computation-storage collaboration demands deeper system-level coordination but offers greater potential to proactively identify new data path challenges and establish novel device interfaces as model architectures and inference infrastructure continue to evolve.
The three approaches are not mutually exclusive. Phison’s strength lies in rapidly packaging its capabilities into deployable solutions using mature controllers, SSD products, and software toolchains; Longsys aims to shift more processing and scheduling to the storage side using SPU and iSA; YinpU–Artemis, on the other hand, defines devices by working backward from the intersection of computational semantics and storage execution. Together, they illustrate that the core innovation in AI SSDs has shifted from “replacing a drive with a faster one for AI” to “redefining the responsibilities among Runtime, memory hierarchy, controller, and Flash.” This is also why AI SSDs have become the pioneering functional SSDs: once software can communicate object types, priorities, lifecycles, and service durations to the device, the SSD gains a foundational capability to support more deterministic functions.
Similar development approaches could readily be applied to the next generation of functional SSDs. Following the Phison path, vendors can integrate encryption, compression, indexing, and memory service modules managed by agent or tenant into SSDs combined with middleware, lowering deployment barriers through software compatibility and delivery tools. Following the Longsys path, compression, hot-cold data scheduling, indexing scans, or data compaction can be executed by the storage processing unit, with policies then orchestrated by upper-layer intelligent schedulers. Following the Yinpuc–ArchiTek path, one could begin by defining requirements for long-term memory, permissions, TTL, retrieval SLOs, and token costs at the agent runtime level, then collaboratively design controller commands, object metadata, firmware queues, and media layouts. The first two paths excel in productization and device-side execution, respectively, while the third path—by simultaneously understanding how computing systems consume data and how storage systems organize it—may offer greater design flexibility when exploring new functionalities that have yet to be standardized.
The significance of this cross-domain background will be further amplified in the Agent era. In the future, many storage functions will no longer be mere on-disk algorithms: encryption requires understanding the permission boundaries between Agents, users, and tasks; compression needs to know when data will be recomputed; indexing must collaborate with Embedding, Retriever, and Model Router; and Agent Memory involves writing, merging, forgetting, versioning, and evidence chains. Focusing solely on the media side risks reducing the problem to bandwidth, capacity, or individual operators; focusing only on the model side may overlook FTL, write amplification, tail latency, power-loss protection, and mass-production constraints. Joint R&D approaches like Yinpu–UnionMemory, which span computing and storage, are uniquely positioned to identify practical interfaces between these two sets of constraints, giving them the potential to develop differentiated technical pathways for next-generation SSDs, AI storage nodes, and Agent data infrastructure.
From a functional SSD perspective, the importance of these solutions lies not only in their ability to scale a model, but in that they establish a channel for software to exchange information with the storage medium. Today, this includes model layers, experts, KV blocks, and prefetch hints; in the future, it can expand to agent identities, memory objects, versions, TTLs, access policies, indexing hints, and latest completion deadlines. Whoever can turn functionality into a stable interface—rather than creating one-off paths tailored for a single model—will have a greater chance of crossing model generations and transforming single-drive revenue into runtime licensing, storage node services, context services, and data services billed by valid tokens.
Similarly, be cautious of conceptual expansion. Automatic compression does not mean all data saves the same amount of space; on-disk indexing does not mean SSDs understand semantics; and Agent Memory does not mean moving a vector database into firmware. Each feature must be proven by end-to-end metrics, including compression ratio and additional latency, recall rate and QPS, P99 and SLO, write amplification and endurance, key isolation and deletion verification, and ultimately, tokens per second and tokens per watt.

Figure 4: Representative enterprise SSDs with AI workload enhancement: YinChen Technologies Dongting N3X and Huawei OceanDisk LC 560. The original figure is retained; product details are available in [19][20].

Figure 5: Exploration of AI SSDs with Inference Participation: Phison aiDAPTIV, Longsys SPU+iSA, and YinpU—ArchiTek AI SSD solutions. Original figure retained; route information from [15][16][17][18].

Figure 6: The role of AI SSD in the LLM inference path and the main technical pathways of vendors outlined in the original manuscript. The original figure is retained; the additional industry extensions introduced in this paper—encryption, compression, indexing, Agent Memory, HBF, and hierarchical storage—are built upon this foundation without altering the existing logic in the diagram.
Future industries may develop simultaneously along three directions.
The first direction is the functionalization of SSDs. General-purpose drives will continue to exist, but capabilities such as security, compression, retrieval, memory, and context will be integrated into products as hardened functions, downloadable programs, or software-defined configurations. The market may ultimately not require five distinct types of SSDs; instead, it is more likely to need a discoverable, composable, and isolatable framework of functions, with consumer, enterprise, and cloud service products selecting different combinations of capabilities.
The second direction is the nodalization of storage. While a single drive addresses local capacity and near-data processing, AI storage nodes integrate multiple drives, networking, object directories, keys, indexes, context services, and observability to own end-to-end SLOs. Edge-side nodes serve individuals and device fleets, while cloud-side nodes serve GPU clusters and multi-agent systems. Competition will shift from “per drive” to “how many effective tokens and retrieval requests each site, rack, or POD can deliver.”
The third direction is the realignment of memory hierarchies. HBM continues to pursue maximum bandwidth, HBF aims to provide larger read-intensive capacity near the package using NAND, DRAM and CXL handle variable states, expansion, and pooling, functional SSDs deliver persistent objects and near-data services, and shared storage preserves global cold data sources. The most valuable future system capability will be enabling routers and runtimes to simultaneously perceive compute, data location, permissions, lifetimes, and media states, coordinating across multiple levels rather than continuously stacking single media types.
This will redefine industry barriers: media manufacturers control capacity, bandwidth, and energy efficiency; controller and firmware vendors determine whether features can be reliably implemented; runtime and agent platforms govern object semantics and scheduling; and OEMs, cloud providers, and system integrators decide how features are integrated into real products. Companies that can transcend these boundaries, establish standardized interfaces, and validate customer outcomes through token economics are more likely to elevate SSDs from mere data containers to the foundational data layer of the Agent era.
Conclusion: Storage will become part of the Agent's capabilities.
The AI Agent expands the value of storage from "preserving the past" to "enabling the next action." Models require weights, reasoning needs context, and agents demand long-term memory and trustworthy evidence. Enterprises also require security, compliance, auditing, and cost control. The closer encryption, compression, indexing, memory maintenance, and context services are to the data, the greater the opportunity to reduce data movement, recalculation, and DRAM usage—but this also demands clear definitions of permissions, semantics, and responsibility boundaries.
Therefore, AI SSD is just the beginning. In the future, more SSDs with specialized functionalities may emerge, potentially leading to standardized programmable functional SSDs and Agent storage nodes. Meanwhile, HBM, HBF, DRAM, CXL, and SSD will not evolve along a simple replacement chain but will instead be redefined based on热度, volatility, sharing scope, and access duration. The real industry opportunity lies in leveraging the strengths of each medium to achieve lower token costs, higher Agent continuity, and more trustworthy data lifecycles.
Reference materials:
[1] NVIDIA, “Scaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueField,” 2026. https://developer.nvidia.com/blog/scaling-agentic-ai-factories-through-extreme-co-design-with-nvidia-bluefield/
[2] NVIDIA, "Introducing NVIDIA BlueField-4-Powered CMX Context Memory Storage Platform for the Next Frontier of AI," 2026. https://developer.nvidia.com/blog/introducing-nvidia-bluefield-4-powered-inference-context-memory-storage-platform-for-the-next-frontier-of-ai/
[3] R. Qin et al., “Mooncake: Trading More Storage for Less Computation—A KVCache-Centric Architecture for Serving LLM Chatbots,” USENIX FAST ’25, 2025. https://www.usenix.org/conference/fast25/presentation/qin
[4] SNIA, "Computational Storage Architecture and Programming Model, Version 1.0," 2022. https://www.snia.org/sites/default/files/technical-work/computational/release/SNIA-Computational-Storage-Architecture-and-Programming-Model-1.0.pdf
[5] SNIA, “Storage Security: Encryption and Key Management,” 2023. https://www.snia.org/sites/default/files/technical-work/whitepapers/SNIA-Encryption-KM-WP-2023-09-05.pdf
[6] NVM Express, “NVM Express Releases Specifications to Unify AI, Cloud, Client, and Enterprise Storage,” 2024. https://nvmexpress.org/nvm-express-releases-nvm-express-specifications-to-unify-ai-cloud-client-and-enterprise-storage/
[7] Samsung Electronics, "Samsung Electronics Develops Second-Generation SmartSSD Computational Storage Drive," 2022. https://news.samsung.com/global/samsung-electronics-develops-second-generation-smartssd-computational-storage-drive-with-upgraded-processing-functionality
[8] KIOXIA, “AiSAQ Achieves 4.8 Billion High-Dimensional Vector Search Database on a Single Server,” 2026. https://americas.kioxia.com/en-us/business/news/2026/ssd-20260316-2.html
[9] Y. Wang et al., "Are We Ready for an Agent-Native Memory System?", arXiv:2606.24775, 2026. https://arxiv.org/abs/2606.24775
[10] P. Chhikara et al., "Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory," arXiv:2504.19413, 2025. https://arxiv.org/abs/2504.19413
[11] SK hynix, “The Next-Generation Memory Architecture in the AI Era? SK hynix Charts the Direction at FMS 2026,” 2026. https://news.skhynix.com/en/fms-2026/
[12] SK hynix, “SK hynix Unveils First HBF Standard Specifications with SanDisk,” 2026. https://news.skhynix.com/en/hbf-at-fms-2026/
[13] M. Ha, E. Kim, and H. Kim, "H3: Hybrid Architecture Using High Bandwidth Memory and High Bandwidth Flash for Cost-Efficient LLM Inference," IEEE Computer Architecture Letters, 2026. https://ieeexplore.ieee.org/document/11371745
[14] Compute Express Link Consortium, “Overcoming the AI Memory Wall: How CXL Memory Pooling Powers Scalable AI Computing,” 2025. https://computeexpresslink.org/blog/overcoming-the-ai-memory-wall-how-cxl-memory-pooling-powers-the-next-leap-in-scalable-ai-computing-4267/
[15] Phison Electronics, "How aiDAPTIV+ Works," official product documentation. https://phisonaidaptiv.com/zh-tw/how-aidaptiv-works/
[16] Longsys, "SPU and iSA," 2026. https://cn.longsys.com/about/news/13353.html
[17] Maxio Technology, "CFMS 2026 | The Value Leap of Storage Controller Chips in the AI Inference Era," 2026. https://www.maxio-tech.com/news/11645/13048.html
[18] Economic Observer, "Yinpu Computing Partners with AMD to Launch Infplane Mini AI Workstation: Hilbert," 2025. https://www.eeo.com.cn/2025/1222/774317.shtml
[19] Yingren Technology, "From N3X to Gen6: How Yingren Technology Built a Domestic AI SSD with Three Key Elements," 2026. https://www.yingren.cn/news/%E4%BB%8En3x%E5%88%B0gen6%EF%BC%9A%E8%8B%B1%E9%9F%A7%E7%A7%91%E6%8A%80%E5%A6%82%E4%BD%95%E7%94%A8%E4%B8%89%E5%A4%A7%E8%A6%81%E7%B4%A0%E6%89%93%E9%80%A0%E5%9B%BD%E4%BA%A7ai-ssd/
[20] Huawei, "Huawei OceanDisk LC 560 SSD Data Sheet," 2025. https://e.huawei.com/en/documents/products/storage/97dc7a1dc98f4d3d90b268db03235cf7
This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation, edited by Solomon.
