- GenStorAIGE AI90 shifts AI memory beyond traditional GPU HBM limitations using SSDs
- The PT200Z SSD supports constant cache updates under demanding inference workloads efficiently
- Eight RTX 5090 GPUs gain dramatically greater effective inference memory capacity
GenStorAIGE has introduced its AI90 inference acceleration platform at WAIC 2026, taking a storage-centric approach to expanding effective AI memory capacity.
Instead of relying solely on high-bandwidth GPU memory, the platform incorporates PCIe Gen5 solid-state drives directly into the memory hierarchy itself.
This allows parts of the Key-Value Cache used by large language models to sit entirely outside of GPU memory.
Latest videos fromTechRadar
A three-tier memory architecture built around SSD offloading
The AI90 combines HBM, system DRAM and SSD into a unified three-layer memory structure to handle inference workloads.
By transparently offloading KV Cache data to SSDs, the platform reduces pressure on GPU memory while supporting significantly larger workloads and longer context windows.
According to GenStorAIGE, this architecture cuts first-token latency from several seconds to sub-second response times in supported configurations.
That represents an improvement of up to 50x, alongside throughput gains of 5.1x and a reduction of approx. 39% in GPU memory usage.
Combined with intelligent peer-to-peer GPU communication, the company states that AI90 can accelerate inference by up to 5.8x on systems running eight Nvidia GeForce RTX 5090 cards.
This multiplier effectively allows an eight-card setup to behave closer to a 46-GPU cluster during persistent inference tasks.
The architecture also supports context windows exceeding 128,000 tokens, enabling much larger document processing and conversation handling without consuming available memory.
The PT200Z SSD handles the intensive write requirements behind the system
To support continuous write workloads generated by constant KV cache updates, GenStorAIGE paired the AI90 with its new PT200Z AI SSD.
Built using pSLC NAND flash and connected via a PCIe Gen5 x4 interface, the drive delivers sequential read speeds that reach 14.8GB/s.
Random read performance hits approximately 3.1 million IOPS, while read latency is only 54 microseconds.
The write event drops even further to 10 microseconds, supporting the fast cache updates that the AI90’s architecture constantly depends on.
Endurance ratings reach up to 100 drive writes per day, a number suitable for persistent enterprise AI workloads with constantly changing cache data.
This design reflects a broader shift across AI infrastructure towards memory tiers as LLMs increasingly outgrow the practical limits of GPU HBM alone.
Integrating extremely fast SSD storage into inference pipelines offers one method to scale context length without requiring additional GPUs or larger HBM configurations.
Whether the performance claims hold up outside of controlled test conditions remains unconfirmed at this stage.
As with most vendor announcements, these performance numbers come directly from GenStorAIGE and still require independent benchmarking across different real-world AI workloads.
Via The Guru of 3D
Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews and opinions in your feeds.



