NVIDIA has introduced two new hardware systems — the Groq 3 LPX and the Vera Rubin NVL72 — designed to accelerate AI inference while cutting costs. Both target the bottleneck of token generation, a key factor in how quickly and cheaply large language models respond. The move has direct implications for agentic AI infrastructure, where performance and price per token are critical.
Inference-Focused Hardware
The Groq 3 LPX and Vera Rubin NVL72 are not for training AI models. They're built for inference, the stage where a trained model generates output. That's a growing demand as AI applications shift from experimental to production, and the cost of running models at scale becomes a dominant factor.
Token generation is the heartbeat of that process. It determines latency, user experience, and the bottom line for AI providers. The two systems are engineered to produce tokens faster, which directly reduces the cost per token. For operators, that means lower expenses for the same level of service.
The Cost-Speed Equation
Inference costs add up quickly when models run round-the-clock. Every token generated consumes compute. Faster generation means less time spent per request, and fewer resources used across thousands of calls. The result is a lower total cost of ownership, which can make a difference for companies deploying AI at scale.
That's where the new hardware enters. By improving token throughput, these systems target the expense side of AI. They're not just about raw performance; they're about making inference more affordable, a shift that could let more organizations run AI without breaking their budgets.
Impact on Agentic AI
Agentic AI — systems that autonomously plan and execute tasks — is a heavy consumer of inference. An agent might make multiple calls to a model in a single workflow, generating tokens at each step. That multiplies cost and latency. The new systems are built to handle that load, with faster token generation that keeps agents responsive and operational costs manageable.
For NVIDIA, this is about positioning at the core of agentic infrastructure. The hardware supports the intensive, sequential reasoning that agents require. Lower cost per token means agents can take more steps before hitting a budget ceiling, potentially expanding what they can accomplish.
The Groq 3 LPX and Vera Rubin NVL72 are now part of NVIDIA's lineup for inference deployments. They're aimed squarely at the growing market for agentic AI, where performance and price are often the deciding factors. How they'll be adopted in real-world systems is the question that will play out in the coming months.



