financetom
Technology
financetom
/
Technology
/
ScaleFlux Introduces AI-Optimized SSD Platform Designed for NVIDIA CMX and KV Cache Offload
News World Market Environment Technology Personal Finance Politics Retail Business Economy Cryptocurrency Forex Stocks Market Commodities
ScaleFlux Introduces AI-Optimized SSD Platform Designed for NVIDIA CMX and KV Cache Offload
Jul 30, 2026 6:16 AM

Platform combines Context-Insight workload intelligence, 7-10+ effective DWPD, and support for more than 200 FDP write streams per drive to improve context memory storage efficiency and economics

MILPITAS, Calif., July 30, 2026 /PRNewswire/ -- ScaleFlux, a leader in advanced Memory controller and storage optimization technology, today announced an AI-optimized SSD platform designed to support the SSD requirements of NVIDIA CMX and other KV-cache-intensive AI inference infrastructure. The platform combines Context-Insight SSD for workload analysis and optimization, 7 to more than 10 effective drive writes per day (DWPD) for KV cache workloads, and support for more than 200 Flexible Data Placement (FDP) write streams per drive for fine-grained lifecycle-aware data placement.

The platform addresses three connected challenges that emerge as AI inference systems use SSDs as a shared context tier beyond GPU HBM and host memory: understanding real workload behavior, separating KV cache blocks with different lifecycles, and sustaining intensive write workloads without excessive SSD capacity inflation or replacement cost.

Long-context inference, shared-prefix reuse, agentic applications, and idle-session retention are rapidly expanding the amount of reusable KV cache and other runtime state that AI systems must maintain. SSDs can provide a cost-effective shared context tier, but this role places new demands on the drive. KV blocks may be written frequently, retained for different periods, reactivated after becoming idle, and invalidated asynchronously across many sessions, workers, or tenants.

ScaleFlux's high-endurance architecture is designed to deliver 7 to more than 10 effective DWPD at 5 years for KV cache workloads, depending on workload characteristics, FDP utilization, and device configuration. Higher effective endurance can reduce the amount of SSD capacity that infrastructure operators must deploy merely to absorb write traffic, allowing more installed capacity to hold useful KV cache and other AI runtime state rather than serving primarily as endurance overhead. This directly addresses the endurance-driven capacity cost (or "endurance tax") associated with high-churn AI inference runtime-state workloads.

The ScaleFlux Context-Insight SSD capabilities are meant to complement the recently-announced NVIDIA CMX Context Memory Storage Platform design and offer AI Factories a high-performance, high-efficiency storage system optimized for KV cache. CMX provides a shared, pod-level context tier for high-speed KV-cache access and reuse, while ScaleFlux helps address the endurance, data-placement, and write-amplification requirements of the SSD tier.

Support for more than 200 FDP write streams per drive enables system software to separate KV data according to expected lifecycle, session or tenant ownership, shared-prefix classification, reuse behavior, or other software-defined categories. Keeping data with similar lifecycles together can reduce garbage-collection movement, lower write amplification, limit interference among data classes, and further improve effective endurance. In preliminary controlled testing, ScaleFlux measured more than a twofold reduction in write amplification using lifecycle-aware FDP placement compared with a baseline placement configuration. Actual results will depend on workload characteristics, lifecycle classification, software integration, and device configuration.

Context-Insight SSD complements these direct endurance and placement capabilities by showing how software-level KV cache policies affect actual SSD behavior. The platform analyzes latency, queue depth, throughput, request-size distribution, data age, write-to-first-read timing, read reuse, NAND write volume, garbage-collection movement, and write amplification. In SSD-only mode, Context-Insight can begin workload characterization without requiring changes to the upper software stack. With software integration, it can correlate high-fidelity SSD telemetry with metadata such as session ID, worker or tenant ID, shared-prefix ID, KV-block ownership, lifecycle state, and key-to-block mapping. This ownership-aware analysis can help identify which sessions, prefixes, tenants, or lifecycle classes are driving SSD latency, endurance consumption, and write amplification.

The combination gives ScaleFlux and its partners a practical path from measurement to production optimization. Context-Insight helps identify and quantify storage-efficiency opportunities, while scalable FDP placement and 7-10+ effective DWPD provide the mechanisms needed to convert those insights into lower write amplification, stronger endurance, and improved AI infrastructure economics.

"AI inference infrastructure needs SSDs that provide more than additional capacity," said Hao Zhong, CEO and Co-Founder of ScaleFlux. "Infrastructure teams need to understand how KV workloads affect the drive, separate data according to lifecycle, and sustain high write rates without deploying excess capacity simply to dilute writes. ScaleFlux brings workload intelligence, scalable FDP placement, and 7-10+ effective DWPD together in one AI-optimized SSD platform."

ScaleFlux is also developing a trace-driven simulator that models KV cache movement across GPU HBM, host memory, and SSD tiers. The simulator generates replayable SSD traces for evaluating placement, eviction, and lifecycle-grouping policies under controlled conditions.

"As AI inference systems extend KV cache beyond GPU Memory and DRAM, understanding the behavior and requirements of the SSD tier becomes increasingly important," said Jason Hardy, Vice President of Storage Technology at NVIDIA. "Our engagement with ScaleFlux is helping characterize how KV cache offload affects storage requirements for latency, endurance, and write amplification, contributing to the broader storage ecosystem around NVIDIA CMX."

ScaleFlux plans to showcase the platform at FMS, including Context-Insight workload analysis, KV metadata correlation, support for more than 200 FDP write streams per drive, lifecycle-aware placement, write-amplification reduction, and high-endurance operation for write-intensive KV cache workloads.

The platform reflects ScaleFlux's broader strategy of enabling SSDs to support increasingly valuable AI runtime state through workload intelligence, software-assisted lifecycle placement, and controller-level endurance optimization.

About ScaleFlux

ScaleFlux is a semiconductor solutions company delivering advanced storage and memory technologies designed to transform data infrastructure in a scalable and sustainable manner. Founded in 2014 and headquartered in Milpitas, California, ScaleFlux develops innovative storage and memory controller technologies for AI, cloud computing, data center, enterprise, and edge applications.

For more information, visit www.scaleflux.com.

View original content to download multimedia:https://www.prnewswire.com/news-releases/scaleflux-introduces-ai-optimized-ssd-platform-designed-for-nvidia-cmx-and-kv-cache-offload-302838473.html

SOURCE ScaleFlux, Inc.

Comments
Welcome to financetom comments! Please keep conversations courteous and on-topic. To fosterproductive and respectful conversations, you may see comments from our Community Managers.
Sign up to post
Sort by
Show More Comments
Related Articles >
What 5 Analyst Ratings Have To Say About Five9
What 5 Analyst Ratings Have To Say About Five9
Aug 1, 2025
5 analysts have shared their evaluations of Five9 ( FIVN ) during the recent three months, expressing a mix of bullish and bearish perspectives. The following table summarizes their recent ratings, shedding light on the changing sentiments within the past 30 days and comparing them to the preceding months. Bullish Somewhat Bullish Indifferent Somewhat Bearish Bearish Total Ratings 2 3...
This CyberArk Software Analyst Is No Longer Bullish; Here Are Top 5 Downgrades For Friday
This CyberArk Software Analyst Is No Longer Bullish; Here Are Top 5 Downgrades For Friday
Aug 1, 2025
Top Wall Street analysts changed their outlook on these top names. For a complete view of all analyst rating changes, including upgrades, downgrades and initiations, please see our analyst ratings page. Baird analyst Shrenik Kothari downgraded the rating for CyberArk Software Ltd ( CYBR ). from Outperform to Neutral and maintained the price target of $460. CyberArk Software ( CYBR...
Analyst Expectations For Monolithic Power Systems's Future
Analyst Expectations For Monolithic Power Systems's Future
Aug 1, 2025
Analysts' ratings for Monolithic Power Systems ( MPWR ) over the last quarter vary from bullish to bearish, as provided by 8 analysts. The table below offers a condensed view of their recent ratings, showcasing the changing sentiments over the past 30 days and comparing them to the preceding months. Bullish Somewhat Bullish Indifferent Somewhat Bearish Bearish Total Ratings 3...
Where Cognex Stands With Analysts
Where Cognex Stands With Analysts
Aug 1, 2025
9 analysts have shared their evaluations of Cognex ( CGNX ) during the recent three months, expressing a mix of bullish and bearish perspectives. The table below summarizes their recent ratings, showcasing the evolving sentiments within the past 30 days and comparing them to the preceding months. Bullish Somewhat Bullish Indifferent Somewhat Bearish Bearish Total Ratings 2 1 4 0...
Copyright 2023-2026 - www.financetom.com All Rights Reserved