Nvidia said at Hot Chips 2026 that the Groq 3 LPX, a specialised inference accelerator, has entered mass production alongside its Vera Rubin platform.

The chip comes out of last December's $20 billion Groq deal, in which Nvidia licensed the startup's technology and hired about 20 of its people, including the founder, rather than buying the company.

The first buyer is the cloud provider Nebius. A full rack holds up to 256 Groq 3 chips and, on Artificial Analysis's measurements, runs Gemma 4 31B at 3,400 tokens a second — roughly twice what Cerebras manages on the same model. The speed comes from memory sitting on the chip itself rather than beside it.

Nvidia is aiming the speed at AI agents and at coding tools that have to answer without long waits. “For folks who are serving tokens, it unlocks the ability to offer premium tiers of service,” said Dion Harris, a senior director at the company.

Jensen Huang said he intends to hand a quarter of the data-centre capacity used for coding workloads to Groq processors. “The rest of my data center is all 100% Vera Rubin,” he said. The two lines are separate down to the fab: Samsung builds the Groq chips, Nvidia's GPUs come off TSMC's.


Cover image: “Nvidia Sign at Taipei, Neihu District, Taiwan Nov 09, 2024 03-04-00 PM” by Thingreenline4546, Wikimedia Commons, CC BY-SA 4.0, resized, re-encoded.