NVIDIA Groq 3 LPX enters full production, targeting agentic AI inference

As seen on the 24/7 Wall St. homepage on August 24, 2026.

PRESS RELEASE NVIDIA

NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI

Full production means revenue timing: the LPX accelerator extends Vera Rubin NVL72 into the latency-sensitive agentic inference market, where Nvidia claims 4x faster agent responsiveness than the nearest alternative platform. Nebius is the first AI cloud deploying it via its Token Factory, with inference cloud Groq lined up next.

Continue ReadingShow less

Full production means NVIDIA can now recognize revenue from Groq 3 LPX, an inference accelerator extending the Vera Rubin platform for agentic systems that need ultrafast token generation.

Groq 3 LPX delivers 4x faster agent responsiveness than the nearest alternative platform, a meaningful competitive wedge in an agentic AI market where user experience depends on how quickly a model can respond and act.

Nebius is the first AI cloud provider to deploy the accelerator, through its Token Factory inference service. Inference cloud Groq is lined up as the next deployment partner, so early commercial distribution is already underway.

The Vera Rubin NVL72 system was already NVIDIA's flagship data-center platform, and Groq 3 LPX extends its reach into latency-sensitive inference workloads. That broadens the addressable market for the Vera Rubin architecture and adds a product tier aimed squarely at real-time agentic applications.

Mentioned: NVDA