๐ฃ Blackwell sets a new inference speed world record โ A single NVIDIA DGX B200 server with eight
#NVIDIABlackwell GPUs can generate over 1,000 tokens per second (TPS) per user on the Llama 4 Maverick model. โกโฑ๏ธ
Additionally, a system with eight Blackwell GPUs can also deliver up to 72,000 tokens/second in a maximum throughput scenario.
๐ Blackwell is the first platform to achieve this model performance, demonstrating how it delivers the best combination of throughput, output speed and accuracy for LLM token generation. ๐๏ธ๐
Check the link in our bio for details