Huawei has introduced a new storage architecture aimed at addressing the rapidly expanding amount of key-value (KV) cache data generated by long-context and multi-turn workloads, which has emerged as a major bottleneck in large-scale AI inference.
The OceanStor M900 Context Memory Storage, unveiled at HUAWEI CONNECT 2026 in Shanghai, is designed to provide hyperscale AI systems with a shared, multi-tier memory space spanning on-chip memory, DRAM and SSD storage.
Designed for AI inference in hyperscale data centers, the product provides SuperPoDs with a fully shared memory space that offers PB-scale capacity and TB/s-level performance.
This marks a shift in AI infrastructure from a compute-centric model to network and storage-centric models. It also marks a broader shift in AI infrastructure as increasingly capable models require more memory to maintain context during inference.
Huawei says models are approaching the 10-trillion-parameter scale, while context windows of more than one million tokens are becoming increasingly common. As a result, KV cache — the intermediate data retained during inference to avoid repeatedly recomputing information — can become a significant consumer of expensive accelerator memory.
As KV cache data generated during inference continues to grow, on-chip memory and DRAM are being pushed beyond their limits in capacity and cost-effectiveness. Rather than keep the entire cache in on-chip memory and DRAM, the M900 extends the memory hierarchy into SSD storage, since SSDs are made up of memory chips anyway. Huawei’s UnifiedBus interconnect is used to create a globally shared KV-cache pool that can distribute data across different memory and storage tiers.
As models grow to trillions of parameters, SuperPoDs are becoming the optimal choice for AI infrastructure. SuperPoDs are sold by NVIDIA and AMD and consist of ultra-dense compute devices packed with GPUs and CPUs. They come with terabytes of memory and can be partitioned up or seen as one giant memory pool.
Huawei says a single M900-based cluster can provide up to 64PB of KV-cache capacity, increasing the amount of cache available per NPU from gigabytes to terabytes. The company argues that the additional capacity can increase cache reuse and improve inference efficiency for long-context workloads.
A key element of the M900 is what Huawei describes as a three-chip integrated architecture combining a CPU, network controller and NAND controller. The design allows an NPU to access SSD-based KV cache through a direct, one-hop path rather than routing the data through a conventional CPU-based storage stack.
Huawei claims the architecture can reduce KV-cache access latency from milliseconds to approximately 60 microseconds, while providing aggregate cluster bandwidth of up to 40TB/s. The company says that, in typical AI programming scenarios, the technology can double inference-cluster token throughput and cut time to first token (TTFT) by half.
The approach effectively treats SSD capacity as an extension of the memory subsystem rather than simply as conventional persistent storage. That distinction is becoming increasingly important as AI inference workloads generate data volumes that cannot economically remain entirely in HBM or DRAM.
The OceanStor M900 represents an attempt to turn storage into an active component of the AI inference pipeline, using high-capacity flash and a specialized interconnect to supplement the limited and expensive memory attached directly to AI accelerators.



