
AI Infrastructure2026-08-25
NVIDIA AI Blog
NVIDIA Vera Rubin NVL72: 30x Efficiency for AI Agents
NVIDIA has unveiled its Vera Rubin NVL72 platform, a new hardware architecture designed to tackle the growing computational demands of AI agents. These agents, which perform complex, multi-step tasks like financial research or comparing peer companies, consume significantly more tokens than simple chat requests—according to OpenRouter data, they use up to 15 times more tokens.
The Vera Rubin NVL72 is engineered to address this challenge head-on. NVIDIA claims the platform delivers up to 30 times more work per watt compared to previous generations. This is a critical leap forward because the token-hungry nature of agentic AI requires not just more compute, but more efficient compute.
By focusing on energy efficiency, the new platform allows organizations to run these sophisticated AI systems without facing prohibitive electricity costs or requiring massive physical infrastructure. The architecture is optimized for the continuous, sequential reasoning that agents require, ensuring that each token generated is done so with maximum efficiency.
For enterprises, this means the barrier to deploying advanced AI agents is lowering. Tasks that were previously too expensive or slow to automate can now be handled in real-time, making AI agents a practical tool for a wider range of business applications. The Vera Rubin NVL72 is not just a hardware upgrade; it is a strategic move to make agentic AI economically viable at scale.