NVIDIA Groq 3 LPX enters full production, targeting agentic AI inference

As seen on the 24/7 Wall St. homepage on August 24, 2026.

PRESS RELEASE NVIDIA

NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI

Full production means revenue timing: the LPX accelerator extends Vera Rubin NVL72 into the latency-sensitive agentic inference market, where Nvidia claims 4x faster agent responsiveness than the nearest alternative platform. Nebius is the first AI cloud deploying it via its Token Factory, with inference cloud Groq lined up next.

Continue ReadingShow less

Full production means NVIDIA can now recognize revenue from Groq 3 LPX, an inference accelerator extending the Vera Rubin platform for agentic systems that need ultrafast token generation.

Groq 3 LPX delivers 4x faster agent responsiveness than the nearest alternative platform, a meaningful competitive wedge in an agentic AI market where user experience depends on how quickly a model can respond and act.

Disclosure

*$149 for two years (or $1.43 per week) is an introductory promotion for new members only. 62% discount based on the current list price of Stock Advisor of $199/year. Membership will renew at the then-current list price at the end of the membership term. Stock Advisor returns are 930% as compared to the S&P 500 returns of 185% as of April 7, 2026.

The same investor newsletter that told subscribers to buy Amazon in 2002, Netflix in 2004, and Nvidia in 2005 still publishes two new stock picks every month. Over 23 years, Motley Fool's Stock Advisor has more than quadrupled the S&P 500. New members get this month's picks, the Top 10 Rankings, and a 30-day money-back guarantee. Click here to unlock their next top stocks while new members are still being accepted.

Nebius is the first AI cloud provider to deploy the accelerator, through its Token Factory inference service. Inference cloud Groq is lined up as the next deployment partner, so early commercial distribution is already underway.

The Vera Rubin NVL72 system was already NVIDIA's flagship data-center platform, and Groq 3 LPX extends its reach into latency-sensitive inference workloads. That broadens the addressable market for the Vera Rubin architecture and adds a product tier aimed squarely at real-time agentic applications.

Mentioned: NVDA