facebook pixel
Inside the NVIDIA Groq 3 LPX compute tray. Each 1U unit packs 8 LPUs with 500MB of SRAM each, 4GB total per unit. The full rack delivers 315 PFLOPS of AI inference compute across 256 chips. The bandwidth story is what stands out: each LPU delivers 150 TB/s of memory bandwidth thanks to SRAM, compared to ~22 TB/s on Rubin’s HBM4. The LPX units interconnect via 4 spines within the rack and use C2C links to connect to adjacent racks. Vera Rubin talks to LPX over Spectrum-X Ethernet. Goal is to offload FFN layers from Rubin GPUs to LPUs to dramatically boost token throughput. Available 2H 2026. #NVIDIA #GTC2026 #Groq

 555

 4

    Suggested Credits
    Tags, Events, and Projects