NVIDIA has put its Groq 3 LPX inference accelerator into full production, commercializing technology acquired through its $20 billion purchase of assets from chip startup Groq in December. Designed to complement the Vera Rubin NVL72 AI platform, the rack targets the decode phase of model serving, where responses are generated token by token, and is aimed at real-time agentic AI, coding and other latency-sensitive workloads. NVIDIA says Groq 3 LPX produced 3,400 output tokens per second in a long-context benchmark using the open-source Gemma 4 31B model, four times the performance of competing platforms. CoreWeave and Nebius are early adopters, while SpaceXAI plans to deploy NVIDIA Vera CPUs in a next-generation architecture spanning data centers and orbital satellites. The systems combine CPUs, GPUs, networking and inference accelerators to reduce cost per token as AI spending shifts toward large-scale reasoning and multi-agent applications. NVIDIA shares were at $210.31 on August 24, down 2.05% over 24 hours, while the company’s market capitalization stood at $5.13 trillion.