Search

Nvidia’s ultra-low-latency AI inference LPX racks hit full production

By: IDCNOVARegion: North America
Nvidia has confirmed that its LPX rack-scale platform for AI inference accelerators has entered full production, marking a significant milestone for the company’s push into ultra-low-latency computing. First unveiled at GTC in March, the platform was developed following Nvidia’s acqui-hire of the eponymous startup, and is designed to address the growing demand for faster, more responsive AI inference at scale.

The LPX rack is liquid-cooled and houses approximately 256 Groq 3 language processing units (LPUs), interconnected through 640 terabits per second of scale-up bandwidth. Inside the rack, the architecture integrates BlueField-4 data processing units, Vera CPU racks, and STX storage servers, all linked using the recently introduced Spectrum-6 Ethernet networking technology. The platform is not positioned as a replacement for Nvidia’s flagship NVL72 offering, but rather as a complementary solution for operators focused on ultra-low-latency AI inference workloads.

During this week’s Hot Chips event in Palo Alto, Nvidia cited industry benchmarking results showing the LPX platform can support 3,400 output tokens per second. The company claims the technology can compress agentic-related tasks from hours to minutes, delivering four times faster responsiveness for AI agents and latency-sensitive applications. Nvidia CEO Jensen Huang highlighted the strategic importance of the platform, stating, “Inference is the growth engine of AI. Nvidia Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency. Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”

The LPX platform is expected to become available later this year. Among the early adopters is neocloud provider Nebius, which plans to integrate the Groq 3 LPX platform into its Token Factory offering. Danila Shtan, Nebius’ chief technology officer, said the move will make “every step of an agent’s loop feel instant.” Groq, the chip and cloud company that has licensed Nvidia’s LPU technology, also confirmed it will be among the first to offer access to the new hardware. Sinclair Schuller, Groq CTO, noted, “Our customers expect Groq to be at the forefront of AI performance, and this platform represents a major advance for next-generation workloads. Together with Dell Technologies, we’re excited to deploy Nvidia Groq 3 LPX and Vera Rubin NVL72 at scale and make this capability broadly available.”

The introduction of the LPX platform underscores a broader industry shift toward specialized, workload-optimized infrastructure for AI inference. As AI agents become more prevalent, the ability to reduce latency and increase token generation throughput is expected to become a key competitive differentiator for cloud providers and enterprises alike. By complementing its flagship NVL72 platform with LPX, Nvidia is positioning itself to serve a wider range of inference use cases, from real-time agent interactions to high-volume token generation, potentially reshaping how AI factories are architected in the coming years.