
AI Infrastructure2026-08-27
NVIDIA AI Blog
With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA is extending its Vera Rubin NVL72 platform to support the next wave of agentic AI, with Groq 3 LPX now in full production. The company made the announcement during a technical briefing, emphasizing that the future of AI inference will be defined by how every layer of the AI factory works together—not by a single breakthrough chip.
Agentic AI systems, which can plan, reason, and execute multi-step tasks, require significantly faster token generation than traditional chatbots. These systems often need to make multiple calls to models, retrieve information from external tools, and iterate on responses in near real-time. The Vera Rubin NVL72 extension is designed specifically to handle this workload, offering low-latency inference that can keep up with the demands of complex agentic workflows.
Groq 3 LPX, now in full production, is a key component of this strategy. It delivers high-speed token generation that reduces the time between user input and final output, making agentic systems feel more responsive and capable. This is especially important for applications like autonomous research, code generation, and customer support automation, where delays can break the user experience.
NVIDIA's approach here is holistic. Instead of optimizing a single component, the company is focusing on the entire inference pipeline—from model loading to memory access to output streaming. The result is a platform that can handle the unpredictable, multi-step nature of agentic AI without sacrificing efficiency.
For developers and enterprises building AI agents, this means they can deploy more sophisticated systems without worrying about infrastructure bottlenecks. As the demand for agentic AI grows, NVIDIA's Vera Rubin NVL72 with Groq 3 LPX is positioned to become a foundational platform for the next era of intelligent automation.