NVIDIA has demonstrated serving the Qwen 3.8 2.4T parameter model with configurable reasoning capabilities on its GB300 NVL72 hardware platform.

  • The system achieves a throughput of over 4,000 tokens per second per GPU using FP8 precision without additional model tuning.
  • Each user receives approximately 350 tokens per second across the 72-GPU configuration.
  • Further performance gains are expected through optimizations such as NVFP4 precision.

This deployment highlights the GB300 NVL72's capacity to handle massive models with high throughput and low latency for end-users.