Search
IDCNOVA

QumulusAI Signs $32 Million Two-Year Deal to Supply NVIDIA Blackwell B300 Capacity for Generative AI Inference

By: IDCNOVARegion: North America
As generative AI workloads surge, the demand for dedicated inference infrastructure has become a critical bottleneck for enterprises scaling image, video, and other compute-intensive models. QumulusAI, a provider of high-performance GPU cloud services, announced a two-year agreement valued at over $32 million to supply NVIDIA Blackwell B300 capacity to an undisclosed AI inference platform provider focused on generative AI applications. The deal includes renewal options, with the capacity expected to come online in the fall of 2026.

The customer’s platform serves developers and enterprises running generative AI models across images, video, and other modalities. Under the agreement, QumulusAI will provide dedicated GPU clusters, ensuring the platform receives committed, high-performance capacity as demand grows. The capacity will be delivered from QumulusAI’s U.S. data center footprint, using what the company describes as a demand-led deployment model. This approach places capacity into available pockets of power across a distributed network of colocation and owned facilities, allowing QumulusAI to bring GPU capacity online within months rather than years.

The agreement adds a two-year commitment of more than $32 million to QumulusAI’s book of business, reflecting a broader industry shift toward dedicated inference infrastructure. “Generative media workloads put real pressure on inference infrastructure; images and video are compute-intensive to serve, and the user experience depends on speed,” said Mike Maniscalco, CEO of QumulusAI. “This agreement reflects a pattern we’re seeing in our own business: inference customers want dedicated, committed capacity they can count on, and our model is built to put that capacity to work quickly.”

The deal underscores the growing importance of inference-stage GPU provisioning as generative AI moves from experimentation to production. By securing long-term, dedicated access to NVIDIA Blackwell B300 clusters, the platform provider gains a competitive edge in delivering low-latency, high-throughput inference for media-heavy AI applications—a segment expected to drive significant data center capacity demand in the coming years.